Ravindra BagaleCourses & study guides मराठी Track your progress

Chapter 17: Scaling

Common Mistakes I See Students Make

These are the scaling mistakes I see after the lab: trying Graviton with an x86 AMI, an email that was never confirmed, and deleting the Auto Scaling group without terminating the instances. Read them once.

  • Changing t3.micro to t3.small and expecting CPU to fall from 90 to a small number. Memory went from 1 GiB to 2 GiB. nproc stayed 2. A 90% CPU problem is the two-instance example, not this type change.
  • Changing the type while the instance is still running, or expecting the public IP of example.com to survive a stop when no Elastic IP is associated. The site looks "down" because SSH is aimed at the old address.
  • Picking t4g.micro for an x86 AMI, then blaming the key pair. lscpu already said x86_64.
  • Launching instance B from a different AMI, or from Amazon Linux with no Nginx, and saying the AMI "did not copy the website". The website is on the AMI you select, not on the instance type.
  • Launching B and never registering it in example-web-tg. B is running. example.com still talks only to A. Horizontal scaling looks like it failed.
  • One datapoint at 95 and a refresh. The worked case that stays OK is 95 then 40. The case that fires is 82 then 91. Two periods.
  • Treating missing data as breaching, so a stopped instance looks like high CPU.
  • The email never arrives because the subscription was not confirmed. Alarm ALARM, topic scale-lab-alerts, subscription still Pending confirmation for student@example.com. Confirm it. Check spam.
  • Two high-CPU alarms. The email is on cpu-high-lab. The policy watches a different alarm the form created. Desired stays 1. Or the mail never comes.
  • Thinking maximum 3 means the alarm launches three instances. The policy adds 1. Desired goes from 1 to 2. Maximum only stops a later jump past 3.
  • Expecting the OK email to scale in. OK is SNS. Desired goes from 2 to 1 because cpu-low-lab and the remove policy said so, or because you set it by hand.
  • Deleting the Auto Scaling group but not terminating the instances, or leaving minimum at 1 and assuming the lab turned itself off.
  • Stopping the only healthy instance before the second one is healthy. Those minutes are downtime.
  • Changing the instance type because the disk is full. df -h at 100% is EBS. Grow the volume.
  • Saying the lab's simple policy went from 2 to 3 by itself. That second add, while CPU stays above 70, is step scaling. The lab uses simple scaling.
  • Leaving desired above 1 overnight, or leaving a larger type running, or keeping a lab AMI and its snapshots.