Ravindra BagaleCourses & study guides मराठी Track your progress

Chapter 17: Scaling

Interview Questions

These scaling questions come up in cloud interviews. In every answer, be clear about scale up and scale out, and about the difference between the alarm and desired capacity.

Q1. What is the difference between vertical and horizontal scaling? Use a real app only as the shape of the problem.
Vertical scaling changes one instance's type. The worked example is example.com on t3.micro (2 vCPU, 1 GiB, CPU about 90%) stopped and started as t3.small (2 vCPU, 2 GiB). There is downtime. Horizontal scaling adds instances of the same size. The worked example is two t3.micro instances behind a load balancer, about 45% CPU each. A viral reel, a YouTube spike, Netflix in the evening, an Amazon sale, Zomato at dinner, or a New Year flood of messages is that horizontal shape: more copies. A bigger database server is the vertical shape. You do not need their private architecture to say that.
Q2. After you stop an instance, change its type and start it, what happens to the public IP and to EBS?
The auto-assigned public IPv4 changes, so a DNS record or an SSH command aimed at the old IP fails. An associated Elastic IP does not change. EBS volumes stay attached, with the same files. The private IP stays. nproc stays 2 when the type goes from t3.micro to t3.small. free -h shows more memory. An x86 AMI still will not boot on Graviton.
Q3. When is manual scaling enough?
When you know the busy hour and you will be there: launch a second instance from the same AMI, register it in the target group, or set desired capacity from 1 to 2 by hand and later back to 1. A demo at 6 pm is manual. A spike at 3 am, or a reel that goes viral while you are away, needs the alarm and the policy.
Q4. What do period and evaluation periods mean? Give the numbers.
The period is 5 minutes, one datapoint. Evaluation is 2 of 2. Average CPU 82 for 10:00–10:05 and 91 for 10:05–10:10 becomes ALARM at the end of the second period. Average 95 then 40 stays OK. INSUFFICIENT_DATA means those datapoints do not exist yet.
Q5. Why can an alarm be in ALARM and still no email arrives?
SNS delivers only to a Confirmed subscription. scale-lab-alerts can be the alarm action, and student@example.com can still be Pending confirmation. The confirm link is in the first SNS mail. The alarm can be right the whole time. Check spam.
Q6. The group is minimum 1, desired 1, maximum 3. The high-CPU policy adds 1. What happens, and what brings desired back to 1?
On ALARM, desired goes from 1 to 2 and one instance launches. Maximum 3 does not launch three instances. It only blocks a later add past 3. Returning to OK sends the SNS email and does not scale in. A second simple policy, on a low-CPU alarm (at or under 30 for 2 periods of 5 minutes), removes 1, so desired goes from 2 to 1. Minimum 1 stops a further remove. Target tracking is the other style: you name a CPU target such as 50 and CloudWatch manages the alarms. This lab keeps the simple policies so the alarms stay visible.
Q7. Why attach a target group to the Auto Scaling group?
So example.com, which points at the load balancer, can reach instances the group launches, and so it stops sending traffic to instances the group terminates. Without the target group, desired still moves from 1 to 2, and visitors still have only the old path.
Q8. Which policy does this lab use, and what do the others do?
The lab uses simple scaling: on 82 then 91, desired goes from 1 to 2, and a low-CPU alarm brings it back toward 1. Manual means you type desired. Step scaling can go from 2 to 3 if CPU stays above 70 after warm-up, and maximum 3 stops 4. Target tracking aims at a level such as 50% CPU. Scheduled scaling sets desired 2 at 08:50 and 1 at 21:00. Predictive scaling forecasts from past load and does not see a brand-new rush.
Q9. What is the order for no downtime?
Keep A healthy behind the load balancer, wait until B is healthy, and only then stop or resize A. A type change of the only instance is downtime. An Elastic IP move is a short blip, not zero.
Q10. How do you tell CPU, RAM, and disk apart?
CPU: utilization about 90% while free -h still shows free memory. Fix with another instance, or a type that adds vCPUs. t3.micro to t3.small does not add a vCPU. RAM: swap is in use, or the kernel kills a process, often because many PHP or Node workers sit on 1 GiB. t3.small is about 2 GiB. Disk: df -h says 100% because of logs, uploads, or database files. Grow the EBS volume (Chapter 13). Changing the instance type does not grow that disk.
Q11. What is the AMI for, in scaling?
The AMI is the copy of the working server: the operating system and the site. A second instance, and every instance the Auto Scaling group launches, must boot from that AMI or the new computer is empty. The launch template names the AMI, the instance type (t3.micro here), the key, and the security group. An x86 AMI does not boot on Graviton.
Q12. What does simple scaling do, and what does step scaling do?
Simple scaling, which this lab uses, adds or removes a fixed number when the alarm enters ALARM, then waits out a cooldown. CPU 82 then 91 moves desired from 1 to 2. It does not add again just because the alarm stays in ALARM. Step scaling looks at how far past the threshold you are, and after a warm-up it can add again if CPU is still above 70: desired 2 to 3. Maximum 3 stops a fourth instance.
Q13. What is target tracking? Give the number.
You name a level, for example keep average CPU near 50%. CloudWatch creates the alarms. One server at 90% is above 50, so the group adds capacity. Two servers near 45% can stay. Near 20% it may scale in, and minimum 1 blocks zero. This lab does not use it, so the alarm you built stays visible.
Q14. What is scheduled scaling? Give the clock.
A scheduled action sets desired at a time you choose. For a 09:00 class in Asia/Kolkata, set desired to 2 at 08:50 and back to 1 at 21:00. Tatkal, a result morning, and dinner fit a clock. A reel at 01:00 does not. Keep a dynamic policy for a surprise.
Q15. What does predictive scaling forecast, and what does it miss?
It forecasts the next hours from past load, and it can raise capacity before a peak you have seen before, such as dinner every evening. It does not know a sale or a reel that has never been in the history. A dynamic policy still has to catch that. This lab leaves predictive scaling off.
Q16. Where does SNS sit in the chain?
Metric, then alarm, then two results of that alarm: SNS emails every confirmed subscription, and the scaling policy changes desired. The email does not launch the instance. A subscription left on Pending confirmation gets nothing. Notify on ALARM and on OK. The OK mail is not scale-in.
Q17. What is instance refresh, in one sentence of order?
The group launches the replacement, waits until it is healthy, and only then terminates the old instance. Visitors always have a healthy target. Stopping the only server first is the outage.
Q18. What happens if you do not scale?
Desired stays 1. CPU climbs. Nginx or Apache can still look running while requests wait, time out, or get a 502 because every worker is busy. If you are asleep, manual scaling does nothing. A sale, tatkal, a result morning, dinner, or a reel does not wait for you to open the console.