Ravindra BagaleCourses & study guides मराठी Track your progress

Chapter 17: Scaling

17.10 No downtime: the order

Why stop-and-resize is downtime

Changing the type of an EC2 instance stops it. Until it is started and healthy again, that instance answers nothing. If it is the only target, example.com is down. There is no zero-downtime option on that one stop.

What actually keeps the site up

Horizontal scaling: the new instance is healthy behind the load balancer before you remove or change the first.

  1. example.com points at the load balancer, not at one public IP.
  2. Instance A is healthy in the target group.
  3. Launch B from the same AMI. Wait until B is healthy. The site still answers.
  4. Only then stop, resize, or deregister A. Visitors who arrive during A's stop still hit B.
  5. Stopping A first, and building B second, is the outage. The gap is the downtime.

How to ship a new update without dropping the site. The old AMI does not contain today's file change. Make a new one, then let the group replace instances one at a time.

  1. Change the site on instance A and test it.
  2. Create a new AMI from A (section 14.2).
  3. Create a new version of launch template example-web-lt that uses that new AMI.
  4. On the Auto Scaling group, start an instance refresh that uses the new template version.
  5. You should see a new instance launch and become healthy.
  6. Only after that should the group terminate an old instance.
  7. You should never see zero healthy targets in the middle.

An Auto Scaling group that replaces an instance launches the new one, waits until the health check passes, and then terminates the old one. Same order. A group cannot change the type of its only instance without a stop. The second instance is the trick.

Instance refresh keeps one server healthy while the next one starts

Moving an Elastic IP from A to B is a short blip, not zero. Open connections on A drop. New ones use B. Do not call that zero downtime.

Cleanup that sets desired to 0 is allowed to take the lab down. That is the end of class, not a change you make in the middle of a sale.

Lab

Chala, write the safe order, one line each.

  1. Point example.com at the load balancer.
  2. Confirm A is healthy.
  3. Launch B from the same AMI.
  4. Wait until B is healthy.
  5. Only then stop or change A.
  6. Now write the wrong order: stop A first.
  7. Mark every minute until B is healthy. Those minutes have zero healthy targets. That is the outage.