Ravindra BagaleCourses & study guides मराठी Track your progress

Chapter 17: Scaling

17.6 Auto Scaling

You can launch instance B by hand once. You cannot sit up for every spike. An Auto Scaling group keeps a count of instances in that range you set, and a scaling policy changes the count when an alarm fires.

Minimum, desired, maximum, with this lab's numbers

Number This lab What it stops you from doing
Minimum 1 The group will not scale in below 1. example.com keeps one server. Desired 0 is something you set at cleanup, not something the low-CPU alarm does
Desired 1, then 2, then back to 1 How many instances should be running now
Maximum 3 A policy that adds 1 can go 1 to 2, and later 2 to 3. It cannot go to 4. If desired is already 3, "add 1" changes nothing

Maximum 3 is not "three instances will appear when the alarm fires". The policy in this lab adds one. The alarm moves desired from 1 to 2. Maximum only caps a later add. Read the activity history if you want to see the sentence AWS wrote: it launched because desired changed, not because the maximum is a wish.

The launch template

A launch template is the recipe for every instance the group creates. It replaces you walking through the launch wizard, which is what section 17.3 did by hand. Create it under EC2, Launch templates.

From the AMI to the launch template, the Auto Scaling group, the alarm, and SNS

Setting Value
Name example-web-lt
AMI The AMI of example.com from section 14.2, the same one as instance B. Not a Graviton AMI
Instance type t3.micro
Key pair The key you already use
Security group Port 80 from the balancer's security group, port 22 from your IP
Subnet Chosen on the group, in the VPC the balancer uses

The template does not run an instance by itself. The group does, when desired is higher than the number of healthy instances.

The group, the target group, and health

Create scale-lab-asg from example-web-lt. Desired 1, minimum 1, maximum 3. The group launches the first instance. Wait until it is InService. You did not press Launch instance. That is the point.

Attach the target group example-web-tg from section 17.3 (or create it now, with the balancer in front, the same way as that section). Instances the group launches are registered. Instances it terminates are removed. example.com, pointed at the balancer, reaches a new instance only because of that registration. Without the target group, desired still changes from 1 to 2. Visitors have no path to the new box.

Health checks. The default is EC2: the instance is running and passing status checks. That does not prove Nginx answers. With a target group attached, set the health check type to ELB so a machine that is "up" but not serving HTTP is replaced. Health check grace period of 300 seconds stops the group from killing an instance that is still booting. A new instance is not healthy in the first second.

Simple scaling, with the numbers, and a word on target tracking

This chapter uses a simple scaling policy so the alarm you built stays the alarm you can point at.

Scale out. Policy name add-one-on-cpu. Type Simple scaling. When cpu-high-lab is in ALARM, add 1 capacity unit. Cooldown 300 seconds. Desired 1 becomes 2. The group launches one t3.micro. Cooldown means a second simple policy will not run again immediately, so one busy stretch does not try to add a third instance in the same minute.

Scale in. A simple policy does not remove capacity when the alarm returns to OK. The OK action is the SNS email from section 17.5. Scale-in is a second alarm and a second policy:

Scale out Scale in
Alarm cpu-high-lab cpu-low-lab
Rule Average CPU greater than 70, 2 of 2 periods of 5 minutes Average CPU less than or equal to 30, 2 of 2 periods of 5 minutes
Policy Add 1 Remove 1
Desired in the worked story 1 → 2 2 → 1
Wall Maximum 3 Minimum 1

Create cpu-low-lab on the group's CPU if the console offers a group metric, or on the instance the group is running, the same way as section 17.4. Notify scale-lab-alerts on ALARM if you want a mail for scale-in too. The policy remove-one-on-cpu watches cpu-low-lab, not cpu-high-lab. One alarm must not be asked to both add and remove.

Target tracking is the other kind of policy, and you should be able to describe it. You would say "keep average CPU near 50". CloudWatch then creates and manages the alarms for you. At 90 on one t3.micro, a 50 target also adds capacity. At about 45 on two instances, it is near that target and can scale in. This lab does not use it, because the alarms would be hidden behind the policy. You already built cpu-high-lab with 82 and 91 on paper. Keep that alarm visible.

If the create-policy form hides "use existing alarm", do not end up with two high-CPU alarms. Move the SNS actions onto the alarm the policy actually uses, and delete the spare.

What happens when the alarm fires, then scale-in

Raise CPU on purpose on the InService instance. SSH in and run:

nproc
yes > /dev/null &

If nproc is more than 1, one yes is not the whole machine:

for i in $(seq $(nproc)); do yes > /dev/null & done

yes writes forever and burns CPU. It is only for this test.

Worked timeline (your clock will differ; the counts will not):

  1. CPU stays high. First 5-minute average, for example 82. Second, for example 91. Both greater than 70.
  2. cpu-high-lab becomes ALARM.
  3. student@example.com gets the ALARM mail, because the subscription is Confirmed.
  4. add-one-on-cpu sets desired from 1 to 2. Activity history on scale-lab-asg shows a launch. A second t3.micro becomes InService and, if the target group is attached, healthy behind the balancer. Maximum is 3, so this step does not launch two extras.
  5. Stop the burn:
pkill yes
  1. CPU falls. After two periods at or under 30, cpu-low-lab becomes ALARM (the low alarm is "in alarm" when CPU is quiet, which sounds backwards and is correct). remove-one-on-cpu sets desired from 2 to 1. The group terminates one instance. Minimum 1 means it stops there. The OK email from cpu-high-lab may arrive as well. That email is not what terminated the instance. The remove policy is.

While you wait, use the Activity tab, not only the inbox. About ten minutes for the high alarm, then about ten minutes of quiet CPU for the low alarm. Cooldown can add a few more minutes between the two policies. That wait is the lab, not a stuck console.

Stop the lab

This is the step that gets skipped.

  1. pkill yes if it is still running. Confirm with top for a moment, or run pkill yes again. It is fine if the process is already gone.
  2. Set minimum to 0 and desired to 0. Wait until the instances are gone. Or delete scale-lab-asg and choose to terminate the instances. If you delete the group and do not terminate, the instances keep running with nobody watching them.
  3. Delete the policies if they were not deleted with the group. Delete cpu-high-lab and cpu-low-lab. Delete the SNS subscription and scale-lab-alerts when you do not want more mail.
  4. Delete the load balancer and example-web-tg if you created them for this lab. A balancer left up is still a front door to instances you think are gone.
  5. If section 17.1 left the original instance on t3.small, change it back to t3.micro.

Do not leave the group running

Maximum 3 means "not more than three", not "it will turn itself off". Minimum 1 keeps one instance even after scale-in to 1. Scale to zero yourself, or delete the group and terminate, before you close the laptop.

Lab

Chala, build the group, fire the alarm, then put it back. One action per line.

  1. Create launch template example-web-lt from the AMI of example.com.
  2. Set the template instance type to t3.micro.
  3. Create Auto Scaling group scale-lab-asg from that template.
  4. Set minimum 1, desired 1, maximum 3.
  5. Attach target group example-web-tg.
  6. Wait until one instance is InService.
  7. Create simple policy add-one-on-cpu that adds 1 when cpu-high-lab is in ALARM.
  8. Create alarm cpu-low-lab: average CPU less than or equal to 30, period 5 minutes, 2 of 2.
  9. Create simple policy remove-one-on-cpu that removes 1 when cpu-low-lab is in ALARM.
  10. SSH to the InService instance.
  11. Run nproc.
  12. Run one yes > /dev/null & per CPU, as in the commands above.
  13. Wait about ten minutes.
  14. You should see alarm state ALARM.
  15. You should see the SNS email.
  16. You should see desired change from 1 to 2.
  17. You should see a second instance InService.
  18. Run pkill yes.
  19. Wait until CPU has been quiet for two periods.
  20. You should see desired change from 2 to 1.
  21. Set minimum to 0.
  22. Set desired to 0.
  23. You should see the instances terminate. Or delete the group and choose terminate.