Ravindra BagaleCourses & study guides मराठी Track your progress

Chapter 18: Load Balancing

18.15 Health checks

Why a dead instance must drop out

Without a health check, the balancer keeps the instance in the list forever. Nginx can be stopped, or the process can crash, and the next user is still sent there. The page hangs. The other healthy instances sit idle. The check exists so a dead EC2 gets no new requests.

What you set

The check is a small request on a timer. These are the values we set for the ordinary web group. They are not a claim that the console opens with these exact numbers. The Application Load Balancer user guide documents common starting values of interval 30 seconds, timeout 5 seconds, healthy threshold 5, unhealthy threshold 2, path /, and success code 200 for an instance target. The API reference has not always agreed on the timeout. Treat the console's pre-filled numbers as may vary, then type the values below.

Setting Value we set
Protocol HTTP
Path /health
Port traffic-port (the same port as the target, 80 for TG-web)
Interval 30 seconds
Timeout 5 seconds
Healthy threshold 2
Unhealthy threshold 2
Success code 200

The arithmetic of those numbers: 2 successful checks, 30 seconds apart, is about 60 seconds. That is the wait when the threshold of 2 is what must pass before the target is marked healthy again. AWS also documents that a target you just registered needs one successful check before it counts as healthy, so the first wait can be closer to one interval, about 30 seconds. Read the target status. Do not invent a third timer.

The app must answer /health from the process itself. A tiny page that returns 200 means "this process is up". Do not make /health call a slow database. If you do, a brief database blip fails the check, the instance is pulled out of service, and the database problem becomes a web outage too.

How Nginx answers /health

On each web instance, Amazon Linux or Ubuntu, Nginx is already installed from section 18.3. Add a location. Reload with sudo service nginx reload.

  1. SSH to the instance.
  2. Open the Nginx site file. On a default lab install the file is /etc/nginx/nginx.conf or a file under /etc/nginx/conf.d/. The path may vary.
  3. Inside the server block, add the location below.
location /health {
    return 200 'ok';
}
  1. Save the file.
  2. Run sudo nginx -t.
  3. You should see a syntax ok line.
  4. Run sudo service nginx reload.
  5. Run curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1/health.
  6. You should see 200.
  7. Repeat on the other instance in the same target group.

What happens after the check

Health checks drop a dead instance ALB EC2-1healthy EC2-2healthy EC2-3unhealthy EC2-4healthy traffic stopped After 2 failed checks, EC2-3 gets no new requests while others stay healthy. Unused / no registered targets → 503. All registered unhealthy → fails open (not always 503).

Four instances: two stay green healthy; one turns red after failed checks and stops receiving new traffic. Empty or unused groups return 503; all unhealthy fails open.

Timeline we set for TG-web t=0check +30scheck ~60shealthy×2 fails×2unhealthy Interval 30s · timeout 5s · healthy threshold 2 · unhealthy threshold 2 · path /health

Interval 30 seconds. About 60 seconds after two successes the target is healthy. Two failures mark it unhealthy.

  1. A healthy target receives new user requests.
  2. An unhealthy target receives none, as long as some other target in that group is healthy.
  3. The Auto Scaling group, if its health check type is ELB, can replace the instance that stays unhealthy. Chapter 17 is that replacement. The balancer does not launch a new EC2 by itself.
  4. If the group has no registered target, or every registered target is unused, the Application Load Balancer returns 503.
  5. If every registered target is unhealthy at the same time, the current health-check page says the Application Load Balancer fails open and still forwards to those targets. Do not memorise "always 503" for that case. Prove it by reading the health column. A 502 is a different error: the balancer reached a target and the answer was bad or the connection closed.
Checks pulse · failed box drops · traffic stays on healthy ALB ok-1 ok-2 fail ok-3 If group empty / unused → 503 All unhealthy → fails open User dots skip the fading failed instance once it is marked unhealthy.

Health checks pulse. A failed instance fades out of rotation while traffic stays on the healthy ones.

How a payment group uses a stricter check

Second example. Checkout should leave a bad box faster than the browsing group.

Setting Browsing group Checkout group
Path /health /health
Interval 30 seconds 10 seconds
Unhealthy threshold 2 2
Time to pull a bad box about 60 seconds about 20 seconds
  1. Set TG-checkout health check interval to 10 seconds.
  2. Set the unhealthy threshold to 2.
  3. Leave the path at /health and the success code at 200.
  4. Two failures, 10 seconds apart, is about 20 seconds.
  5. You should see the checkout target go unhealthy in about 20 seconds, while a browsing target with interval 30 is still on its first gap.

The mistake that marks every target unhealthy

If the path is / and the app returns 302 to a login page, the success code 200 never arrives. Every target looks unhealthy. The fix is a path that returns 200 with no login. /health with return 200 'ok' is that path.

  1. Open the target group health settings.
  2. Read the path. If it is /, that is the suspect.
  3. From the instance, run curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1/.
  4. If you see 302, that is why the check fails.
  5. Change the health check path to /health.
  6. Confirm /health returns 200.
  7. Wait for the healthy threshold.
  8. You should see the targets become healthy without turning the login page off.

Ravindra Bagale's Tip

Many students set the health check path to /. The app returns 302 to the login page, so every target looks unhealthy. Use /health, which returns 200, with no login.