Chapter 17: Scaling
Let's go. In Chapter 14 we made one server bigger: stop, change the instance type, start. Today we look at the whole picture of scaling. Making one server bigger, adding more servers, a CloudWatch alarm, an SNS email, and an Auto Scaling group that changes desired capacity by itself. Don't worry, you will get it.
चला मित्रांनो, Chapter 14 मध्ये आपण एक server मोठा केला: stop, instance type बदल, start. आज scaling चा पूरा picture बघूया. एक server मोठा करणे, जास्त servers लावणे, CloudWatch alarm, SNS ची email, आणि Auto Scaling group जे desired capacity स्वतः बदलते. घबरू नका, कळेल तुम्हाला.
चलो दोस्तों, Chapter 14 में हमने एक server बड़ा किया: stop, instance type बदलो, start. आज scaling की पूरी picture देखें. एक server बड़ा करना, ज्यादा servers लगाना, CloudWatch alarm, SNS की email, और Auto Scaling group जो desired capacity खुद बदलता है. घबराओ मत, समझ आएगा.
What you will learn in this chapter
- Why one server runs out of CPU or memory, and when a bigger type is the wrong fix
- Vertical scaling on example.com:
t3.microat 90% CPU, changed tot3.small - Horizontal scaling: two
t3.microinstances behind one load balancer - Manual scaling: a second instance from the same AMI, and desired capacity by hand
- A CloudWatch alarm with the numbers written out: CPU greater than 70 for 2 periods of 5 minutes
- SNS: create a topic, confirm the email, notify on ALARM and on OK
- An Auto Scaling group with minimum 1, desired 1, maximum 3: the alarm moves desired from 1 to 2, then scale-in brings it back to 1
- A sudden rush in shopping, payments, video, food, tickets, exams, and a viral reel, and what the alarm does while you sleep
- When the limit is CPU, RAM, or the disk
- How Nginx and Apache use workers, and the order that avoids downtime
- Manual, simple, step, target tracking, scheduled, and predictive scaling
- How to put the lab back when class ends, with no prices
- Public incidents where a real site slowed under demand
This chapter uses the EC2 instance from Chapter 5, the type change from section 14.4, and the AMI from section 14.2. The site in every example is example.com. Keep the console in Asia Pacific (Mumbai) ap-south-1 and do every step yourself. When the lab is finished, do not leave extra instances running.
People already know what a traffic spike feels like, even if they have never opened the AWS console.
| What you have seen | The honest lesson | What this chapter will not invent |
|---|---|---|
| An Instagram reel suddenly has a huge number of views | That is a traffic spike. The app needs more copies of the servers that answer viewers. That is horizontal scaling | Instagram's private design. We do not know which products they run, and we will not pretend to |
| WhatsApp on New Year's Eve, when everyone sends "Happy New Year" at once | A flood of messages needs more machines accepting work. One bigger machine would be vertical scaling | WhatsApp's real servers, databases, or how a message is stored |
| A YouTube video that millions of people open in one hour | More viewers means more requests to answer. More servers is horizontal. One bigger server is vertical | YouTube's internal video pipeline |
| Netflix in the evening, when many people press play together | The rush is horizontal: more servers for more viewers at the same time | Netflix's real architecture |
| Amazon on a sale day | The website in front gets a spike of clicks (horizontal). Making one database server larger, because that one box ran out of memory, is vertical | Amazon's real catalog, or how many servers they use |
| Zomato around dinner, when orders jump | More app servers for the dinner rush is horizontal. A bigger database server for one database is vertical | Zomato's real kitchen, payments, or cloud design |
example.com in this lab is a small teaching site, not those apps. The shape of the problem is the same. One machine has a limit. Either you make that machine bigger, or you add another machine and put something in front so users still open one name: example.com.
Visitors type example.com
|
v
Load balancer (one DNS name)
|
+--> t3.micro InService about 45% CPU
|
+--> t3.micro InService about 45% CPU
Before scale-out there was one t3.micro at about 90% CPU.
Chain when CPU stays high:
metric crosses 70 for 2 periods of 5 minutes
--> alarm state ALARM
--> SNS email (only if the subscription is Confirmed)
--> scaling policy sets desired 1 --> 2
--> group launches the second instance
Later, when CPU stays low:
--> scale-in sets desired 2 --> 1
Minimum is 1, so the group does not go to zero.
Maximum is 3, so a later add could reach 3, never 4.
Concepts in this chapter
- 17.1Vertical scaling (scale up)
- 17.2Horizontal scaling (scale out)
- 17.3Manual scaling
- 17.4CloudWatch alarm
- 17.5SNS: email when the alarm changes
- 17.6Auto Scaling
- 17.7What happens when traffic increases
- 17.8CPU, RAM, or storage
- 17.9How Nginx and Apache use CPU and RAM
- 17.10No downtime: the order
- 17.11Every Auto Scaling policy
- 17.12When class ends
- 17.13Public reports, not our clients