18.15 Health checks
Why a dead instance must drop out
Without a health check, the balancer keeps the instance in the list forever. Nginx can be stopped, or the process can crash, and the next user is still sent there. The page hangs. The other healthy instances sit idle. The check exists so a dead EC2 gets no new requests.
What you set
The check is a small request on a timer. These are the values we set for the ordinary web group. They are not a claim that the console opens with these exact numbers. The Application Load Balancer user guide documents common starting values of interval 30 seconds, timeout 5 seconds, healthy threshold 5, unhealthy threshold 2, path /, and success code 200 for an instance target. The API reference has not always agreed on the timeout. Treat the console's pre-filled numbers as may vary, then type the values below.
| Setting | Value we set |
|---|---|
| Protocol | HTTP |
| Path | /health |
| Port | traffic-port (the same port as the target, 80 for TG-web) |
| Interval | 30 seconds |
| Timeout | 5 seconds |
| Healthy threshold | 2 |
| Unhealthy threshold | 2 |
| Success code | 200 |
The arithmetic of those numbers: 2 successful checks, 30 seconds apart, is about 60 seconds. That is the wait when the threshold of 2 is what must pass before the target is marked healthy again. AWS also documents that a target you just registered needs one successful check before it counts as healthy, so the first wait can be closer to one interval, about 30 seconds. Read the target status. Do not invent a third timer.
The app must answer /health from the process itself. A tiny page that returns 200 means "this process is up". Do not make /health call a slow database. If you do, a brief database blip fails the check, the instance is pulled out of service, and the database problem becomes a web outage too.
How Nginx answers /health
On each web instance, Amazon Linux or Ubuntu, Nginx is already installed from section 18.3. Add a location. Reload with sudo service nginx reload.
- SSH to the instance.
- Open the Nginx site file. On a default lab install the file is
/etc/nginx/nginx.confor a file under/etc/nginx/conf.d/. The path may vary. - Inside the server block, add the location below.
location /health {
return 200 'ok';
}
- Save the file.
- Run
sudo nginx -t. - You should see a syntax ok line.
- Run
sudo service nginx reload. - Run
curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1/health. - You should see
200. - Repeat on the other instance in the same target group.
What happens after the check
Four instances: two stay green healthy; one turns red after failed checks and stops receiving new traffic. Empty or unused groups return 503; all unhealthy fails open.
Interval 30 seconds. About 60 seconds after two successes the target is healthy. Two failures mark it unhealthy.
- A healthy target receives new user requests.
- An unhealthy target receives none, as long as some other target in that group is healthy.
- The Auto Scaling group, if its health check type is ELB, can replace the instance that stays unhealthy. Chapter 17 is that replacement. The balancer does not launch a new EC2 by itself.
- If the group has no registered target, or every registered target is unused, the Application Load Balancer returns 503.
- If every registered target is unhealthy at the same time, the current health-check page says the Application Load Balancer fails open and still forwards to those targets. Do not memorise "always 503" for that case. Prove it by reading the health column. A 502 is a different error: the balancer reached a target and the answer was bad or the connection closed.
Health checks pulse. A failed instance fades out of rotation while traffic stays on the healthy ones.
How a payment group uses a stricter check
Second example. Checkout should leave a bad box faster than the browsing group.
| Setting | Browsing group | Checkout group |
|---|---|---|
| Path | /health |
/health |
| Interval | 30 seconds | 10 seconds |
| Unhealthy threshold | 2 | 2 |
| Time to pull a bad box | about 60 seconds | about 20 seconds |
- Set
TG-checkouthealth check interval to 10 seconds. - Set the unhealthy threshold to 2.
- Leave the path at
/healthand the success code at 200. - Two failures, 10 seconds apart, is about 20 seconds.
- You should see the checkout target go unhealthy in about 20 seconds, while a browsing target with interval 30 is still on its first gap.
The mistake that marks every target unhealthy
If the path is / and the app returns 302 to a login page, the success code 200 never arrives. Every target looks unhealthy. The fix is a path that returns 200 with no login. /health with return 200 'ok' is that path.
- Open the target group health settings.
- Read the path. If it is
/, that is the suspect. - From the instance, run
curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1/. - If you see
302, that is why the check fails. - Change the health check path to
/health. - Confirm
/healthreturns200. - Wait for the healthy threshold.
- You should see the targets become healthy without turning the login page off.
Ravindra Bagale's Tip
Many students set the health check path to /. The app returns 302 to the login page, so every target looks unhealthy. Use /health, which returns 200, with no login.
Ravindra Bagale's Tip – मराठी
हेल्थ चेक path / ठेवला, आणि अॅप लॉगिनसाठी 302 देते, तर सगळे targets unhealthy दिसतात. /health ठेवा, जो 200 देतो, लॉगिनशिवाय.
Ravindra Bagale's Tip – हिंदी
हेल्थ चेक path / रखा, और ऐप लॉगिन के लिए 302 देता है, तो सारे targets unhealthy दिखते हैं. /health रखो, जो 200 दे, लॉगिन के बिना.