Ravindra BagaleCourses & study guides मराठी Track your progress

Chapter 17: Scaling

Let's go. In Chapter 14 we made one server bigger: stop, change the instance type, start. Today we look at the whole picture of scaling. Making one server bigger, adding more servers, a CloudWatch alarm, an SNS email, and an Auto Scaling group that changes desired capacity by itself. Don't worry, you will get it.

What you will learn in this chapter

  • Why one server runs out of CPU or memory, and when a bigger type is the wrong fix
  • Vertical scaling on example.com: t3.micro at 90% CPU, changed to t3.small
  • Horizontal scaling: two t3.micro instances behind one load balancer
  • Manual scaling: a second instance from the same AMI, and desired capacity by hand
  • A CloudWatch alarm with the numbers written out: CPU greater than 70 for 2 periods of 5 minutes
  • SNS: create a topic, confirm the email, notify on ALARM and on OK
  • An Auto Scaling group with minimum 1, desired 1, maximum 3: the alarm moves desired from 1 to 2, then scale-in brings it back to 1
  • A sudden rush in shopping, payments, video, food, tickets, exams, and a viral reel, and what the alarm does while you sleep
  • When the limit is CPU, RAM, or the disk
  • How Nginx and Apache use workers, and the order that avoids downtime
  • Manual, simple, step, target tracking, scheduled, and predictive scaling
  • How to put the lab back when class ends, with no prices
  • Public incidents where a real site slowed under demand

This chapter uses the EC2 instance from Chapter 5, the type change from section 14.4, and the AMI from section 14.2. The site in every example is example.com. Keep the console in Asia Pacific (Mumbai) ap-south-1 and do every step yourself. When the lab is finished, do not leave extra instances running.

People already know what a traffic spike feels like, even if they have never opened the AWS console.

What you have seen The honest lesson What this chapter will not invent
An Instagram reel suddenly has a huge number of views That is a traffic spike. The app needs more copies of the servers that answer viewers. That is horizontal scaling Instagram's private design. We do not know which products they run, and we will not pretend to
WhatsApp on New Year's Eve, when everyone sends "Happy New Year" at once A flood of messages needs more machines accepting work. One bigger machine would be vertical scaling WhatsApp's real servers, databases, or how a message is stored
A YouTube video that millions of people open in one hour More viewers means more requests to answer. More servers is horizontal. One bigger server is vertical YouTube's internal video pipeline
Netflix in the evening, when many people press play together The rush is horizontal: more servers for more viewers at the same time Netflix's real architecture
Amazon on a sale day The website in front gets a spike of clicks (horizontal). Making one database server larger, because that one box ran out of memory, is vertical Amazon's real catalog, or how many servers they use
Zomato around dinner, when orders jump More app servers for the dinner rush is horizontal. A bigger database server for one database is vertical Zomato's real kitchen, payments, or cloud design

example.com in this lab is a small teaching site, not those apps. The shape of the problem is the same. One machine has a limit. Either you make that machine bigger, or you add another machine and put something in front so users still open one name: example.com.

  Visitors type example.com
            |
            v
  Load balancer (one DNS name)
            |
            +--> t3.micro   InService   about 45% CPU
            |
            +--> t3.micro   InService   about 45% CPU

  Before scale-out there was one t3.micro at about 90% CPU.

  Chain when CPU stays high:
  metric crosses 70 for 2 periods of 5 minutes
        --> alarm state ALARM
              --> SNS email (only if the subscription is Confirmed)
              --> scaling policy sets desired 1 --> 2
                    --> group launches the second instance
  Later, when CPU stays low:
        --> scale-in sets desired 2 --> 1
  Minimum is 1, so the group does not go to zero.
  Maximum is 3, so a later add could reach 3, never 4.

Concepts in this chapter

  1. 17.1Vertical scaling (scale up)
  2. 17.2Horizontal scaling (scale out)
  3. 17.3Manual scaling
  4. 17.4CloudWatch alarm
  5. 17.5SNS: email when the alarm changes
  6. 17.6Auto Scaling
  7. 17.7What happens when traffic increases
  8. 17.8CPU, RAM, or storage
  9. 17.9How Nginx and Apache use CPU and RAM
  10. 17.10No downtime: the order
  11. 17.11Every Auto Scaling policy
  12. 17.12When class ends
  13. 17.13Public reports, not our clients

Review and practice

  1. ★Common Mistakes I See Students Make
  2. ★Interview Questions
  3. ★Practice Questions and Lab Exercises