Ravindra BagaleCourses & study guides मराठी Track your progress

Chapter 17: Scaling

Practice Questions and Lab Exercises

Now do this lab yourself. When you finish, set desired to zero and terminate the instances.

  1. In one sentence each, define scale up, scale down, scale out and scale in. For one of them, name a famous app only as the shape of a spike (reel, messages, video, evening viewing, sale day, or dinner orders). Do not describe that company's private design.
  2. example.com is a t3.micro at about 90% CPU. You change it to t3.small. What does nproc do, what does free -h do, and why might CPU stay high?
  3. Why will that same instance not start as t4g.small?
  4. Draw two t3.micro instances behind a load balancer. Label the CPU you expect if one instance was at 90 and the visitors split in half. Where does example.com point?
  5. Write the two CloudWatch tables: 82 then 91, and 95 then 40. Which one is ALARM, which one stays OK, and about how many minutes did the first one take?
  6. The subscription is Pending confirmation. The alarm is ALARM. Why is the inbox empty? What do you click?
  7. Minimum 1, desired 1, maximum 3, policy "add 1". The alarm fires once. What is desired afterwards? What is desired after the low-CPU policy removes 1? Why not 0, and why not 4?
  8. Lab. Do the numbered labs in sections 17.1 through 17.12 in that order. The short list is:

    1. Record nproc and free -h on t3.micro.
    2. Change to t3.small, record them again, then change back.
    3. Put a second instance from the same AMI behind the load balancer and see it healthy.
    4. Confirm the SNS email before you trust the alarm.
    5. Create the group at minimum 1, desired 1, maximum 3.
    6. Run yes until desired is 2 and the email has arrived.
    7. Run pkill yes and wait until desired is 1.
    8. Set minimum and desired to 0, and delete the alarms, the subscription, the topic, and the balancer.

Understood? Scaling means making one server bigger or adding more servers. When the alarm crosses the threshold, SNS sends the email, and the policy changes desired. Now repeat the lab without looking at the notes.

Quick Revision

  • A viral reel, a burst of messages, a popular video, an evening of streaming, a sale day, or dinner orders is a traffic spike: more servers (horizontal). One bigger database server is vertical. example.com is the lab, not those companies' real designs.
  • Scale up worked example: stop t3.micro, change to t3.small, start. Downtime. nproc stays 2. free -h goes from about 1 GiB to about 2 GiB. CPU at 90% may stay high. Public IP changes unless an Elastic IP is associated. EBS and the private IP stay. Do not move x86 to Graviton. Change the type back.
  • Scale out worked example: two t3.micro instances behind a load balancer, about 45% CPU each, example.com pointed at the balancer. Each instance has its own disk. The app must survive a request landing on either one.
  • Manual: launch instance B from the same AMI, register it in example-web-tg, or set desired by hand. Enough when you are there. Not enough at 3 am.
  • CloudWatch: average CPUUtilization, period 5 minutes, greater than 70, 2 of 2. 82 then 91 becomes ALARM after about ten minutes. 95 then 40 stays OK. INSUFFICIENT_DATA means the datapoints are not there.
  • SNS: standard topic scale-lab-alerts, email student@example.com, status must be Confirmed or no mail is sent. Notify on ALARM and on OK.
  • Chain: metric, alarm, SNS email, simple scaling policy, desired capacity. The email does not launch the instance.
  • Auto Scaling worked example: minimum 1, desired 1, maximum 3. High alarm adds 1, desired 1 to 2. Low alarm (at or under 30, 2 of 2) removes 1, desired 2 to 1. Maximum does not mean three instances appear at once.
  • A rush (a sale on Flipkart or Amazon, Paytm after salary, a new Netflix title, Zomato or Swiggy at dinner, IRCTC tatkal, exam result day, an Instagram reel) with you asleep leaves desired at 1. The alarm emails you and simple scaling goes 1 to 2. Step scaling can go 2 to 3 if CPU stays above 70. Maximum is 3. Scale-in returns toward 1.
  • CPU, RAM, and disk are different. 90% CPU with free memory is not a full disk. A full disk is EBS, not the instance type. Swap and OOM are RAM.
  • No downtime: B healthy before you stop A. Nginx or Apache still running can mean a queue or a 502. Check with sudo service nginx status or sudo service httpd status.
  • Policies: the lab is simple scaling. Also know manual, step, target 50%, scheduled 09:00, and predictive from past load. After class, desired 0, original type, delete unused AMIs and snapshots.
  • Public reports only: Ticketmaster in November 2022, Pokémon GO in July 2016, HealthCare.gov in October 2013. Not our clients. Do not invent their design.
  • Clean up: pkill yes, minimum 0 and desired 0, or delete the group and terminate. Delete the balancer. Do not leave the instances running.