Ravindra BagaleCourses & study guides मराठी Track your progress

Chapter 17: Scaling

17.1 Vertical scaling (scale up)

Why CPU or memory runs out

example.com is one EC2 instance. Nginx answers the browser. PHP or Node does the work. The operating system needs memory too. A t3.micro has 2 vCPUs and 1 GiB of memory. That is a real, small computer, not a picture of a computer.

Two different things get "full", and students mix them up.

CPU. The processors are busy. CloudWatch CPUUtilization near 90 means, across the last few minutes, the CPUs were busy about 90 percent of the time. Pages feel slow. top shows a process stuck at the top. A viral reel is this feeling at a huge size: too many requests for the CPUs you have. On example.com the same picture is one student lab, not Instagram.

Memory. free -h shows almost no memory available. The kernel starts using swap, and the machine feels frozen even when you expected "only" a web page. A bigger database server is often this problem: one database, more RAM, same job. That is vertical scaling. Amazon's sale day is not "make one web server huge". A single database box that ran out of RAM can be.

An instance type is AWS's name for the size of the virtual machine: how many vCPUs, how much memory, and which CPU architecture. t3.micro and t3.small are two sizes in the same family. You are not installing a new operating system. You are asking AWS to run the same disks on a larger size.

Vertical scaling (scale up) means one server stays one server, and you pick a larger type. Scale down is the same move back to a smaller type. Section 14.4 already changed a type once. This section is that move told as a worked example, with the numbers you should write in your notebook.

Worked example: t3.micro at 90% CPU, changed to t3.small

example.com runs on one instance.

Before After the type change
Name visitors use example.com example.com
Instance type t3.micro t3.small
vCPUs 2 2. nproc still prints 2
Memory 1 GiB 2 GiB. free -h about doubles
CloudWatch CPU about 90% Not guaranteed to fall. Same 2 vCPUs can still sit near 90% if the work is pure CPU
Public IPv4 for example 3.7.1.10 (yours will differ) A new address, unless an Elastic IP is associated
Private IP for example 172.31.10.20 Same private IP
Disk EBS root volume, the site files Same volume, same files

Read that table before you click anything. Students expect nproc to jump. It will not. t3.micro and t3.small both have 2 vCPUs. What you bought with the larger type is memory, 1 GiB to 2 GiB. If example.com was slow because it was out of RAM and swapping, the page gets better. If it was slow because the CPUs were at 90% doing real work, t3.small does not add a CPU. The honest next step is section 17.2, two machines, not a still-bigger type. Do the type change anyway so you have seen stop, change, start with your own eyes. Then decide.

What an instance type change actually does

The instance must be stopped. AWS cannot slide a running machine onto a different size while Nginx is answering example.com. Stop means downtime. For those minutes, example.com does not answer. Do not practice this on the only server a real class is using at that moment. In the lab it is your instance, and a few minutes of downtime is the lesson.

What changes

  • vCPU and memory, to whatever the new type has. Here, memory changes and vCPU count does not.
  • The auto-assigned public IPv4 address. After start, the old address is gone. SSH to the new one. If example.com was a DNS record pointed at that public IP, the name breaks until you edit DNS. An Elastic IP from section 14.1 stays with the instance, so the public address does not change.
  • Which physical host the instance runs on.

What stays

  • EBS volumes and the bytes on them, including the root volume. Your Nginx config, your HTML, your home directory: still there. This course's instances use EBS. Instance-store disks, if a type has them, are wiped by a stop. You do not have those in this lab.
  • The private IP, the instance ID, the security groups, the key pair, and the IAM role.
  • The software you installed. You do not reinstall Linux.

Do the change, and verify it

  1. SSH to the t3.micro.
  2. Run nproc.
  3. Run free -h.
  4. Run lscpu.
  5. Write those three results in a notebook.
  6. In the EC2 console, write down the public IPv4 and the private IP.
nproc
free -h
lscpu

nproc is how many CPUs Linux sees. Expect 2. free -h is memory. The total line is about 1 GiB on a micro (some of it is already used by Linux). lscpu prints the architecture. On this instance it is x86_64, not Arm. That line matters in a moment.

  1. In the EC2 console, select the instance.
  2. Open Instance state.
  3. Choose Stop instance.
  4. Wait until the state is Stopped. example.com does not answer during this wait. Button names may vary.
  5. Open Actions.
  6. Open Instance settings.
  7. Choose Change instance type.
  8. Select t3.small. That is one step up: about 1 GiB of memory becomes about 2 GiB. Do not pick a huge type.
  9. Apply. If the console refuses the type, read the message. Do not pick a Graviton type.
  10. Open Instance state.
  11. Choose Start instance.
  12. Wait until the state is Running.
  13. Copy the new public IP if you have no Elastic IP.
  14. Check the private IP. It should be the one you wrote down.
  15. SSH to the new public IP.
  16. Run nproc again. You should see 2.
  17. Run free -h again. You should see about 2 GiB.
  18. Run lscpu again. You should still see x86_64.
  19. Run sudo service nginx status (or sudo service httpd status if this box is Apache). You should see the service running.
sudo service nginx status

If this instance is Apache, use sudo service httpd status. The service should be running because it was enabled at boot. You did not reinstall it. The disk came back with the unit file on it.

What you should see. nproc still 2. free -h total about 2 GiB. lscpu still x86_64. The private IP unchanged. The public IP changed, unless an Elastic IP is associated. example.com works again only if DNS or your browser is aimed at the address that exists now.

Pause for a minute. While the type changes, the website stops. The public IP changes. The data on EBS stays.

x86 and Graviton

t3 and m5 are x86. t4g and m7g are Graviton (Arm). The AMI is built for one architecture. lscpu said x86_64. Changing this instance to t4g.small does not boot. The key pair is not the problem. The CPU family is the problem. If you want Graviton, launch a new instance from an Arm AMI and move the site. This lab is not that experiment.

When not to scale up

  • The CPUs are the limit and the next type does not add vCPUs. t3.micro at 90% to t3.small is this case. Memory doubles. CPU count does not. Two t3.micro instances (section 17.2) are the lesson that matches a 90% CPU graph.
  • You cannot take downtime. Stop takes example.com off the air. A dinner rush on a food-delivery app, or a reel that is viral right now, is a bad moment to stop the only server.
  • One machine is still one failure. A bigger type dies as one box. Instagram's reel does not stay online because somebody picked a larger size. It stays online because more than one server can answer.
  • The slowness is not this EC2 instance. A slow SQL query on RDS (Chapter 16) stays slow if you only enlarge the web server. Look at CPU and memory here before you change the type. free -h and CloudWatch tell you which resource ran out.
  • You leave the larger type running after the lab. Change it back: stop, t3.micro, start. A short test is not a new permanent size.

Ravindra Bagale's Tip

Many students leave a bigger instance type in place and never think about adding a second server. Look: if the website lives on only one machine, the site is down when that machine is down. Two smaller instances behind a load balancer are the better fit most of the time. Pay attention.

Lab

Chala, put the type back. One action per line.

  1. Stop the instance again.
  2. Change the type from t3.small back to t3.micro.
  3. Start the instance.
  4. SSH in with the public IP you see now.
  5. Run free -h. You should see about 1 GiB again.
  6. Do not leave t3.small running.