17.4 CloudWatch alarm
The second instance in section 17.3 appeared because you noticed. Amazon CloudWatch is how AWS notices. It stores metrics. A metric is a number, measured again and again: CPU percent, not a paragraph. An alarm watches one metric and changes state when your rule has been true for long enough.
The metric in this lab is CPUUtilization of an EC2 instance. Namespace AWS/EC2. Dimension InstanceId, so the number belongs to one instance, not to "CPU somewhere in the account". The value is a percent from 0 to 100. Basic monitoring publishes a point about every 5 minutes. Pick a period of 5 minutes so your alarm and the metric are talking about the same width of time. A 1-minute period does not fill itself in.
The five words, with this lab's numbers
| Piece | This lab | Plain meaning |
|---|---|---|
| Statistic | Average | One number for the whole period, not a single spike inside it |
| Period | 5 minutes | How wide one datapoint is |
| Threshold | Greater than 70 | The line. At 70 the rule is not yet true if you chose "greater than". This lab uses greater than 70 |
| Evaluation periods | 2 | How many periods CloudWatch looks at |
| Datapoints to alarm | 2 out of 2 | Both periods must be over 70. One busy period is not enough |
Worked example: the numbers written out
Imagine the clock on a t3.micro that serves example.com. You do not need these exact minutes in your account. You need to be able to say why the state changed.
Case 1. The alarm fires.
| Period | Average CPUUtilization | Over 70? |
|---|---|---|
| 10:00 to 10:05 | 82 | Yes |
| 10:05 to 10:10 | 91 | Yes |
At about 10:10 CloudWatch has 2 datapoints out of the 2 it required. The state becomes ALARM. One period at 82 was not enough. The second period is what made the rule true. From the moment CPU actually went high, you waited about ten minutes, not ten seconds. That wait is the period times the evaluation count. Anyone who refreshes the alarm after one minute and says "CloudWatch is broken" skipped this table.
Case 2. The alarm does not fire.
| Period | Average CPUUtilization | Over 70? |
|---|---|---|
| 10:00 to 10:05 | 95 | Yes |
| 10:05 to 10:10 | 40 | No |
1 out of 2 is not 2 out of 2. The state stays OK. A one-period spike is exactly what evaluation periods are for. A reel, a sale, or dinner has to stay busy, not blip.
Case 3. There is no number yet.
You created the alarm two minutes ago, or you stopped the instance. CloudWatch does not have those two periods. The state is INSUFFICIENT_DATA. That is not ALARM and it is not a failure. A stopped instance also stops publishing CPU. Treat missing data as missing. If you treat missing data as bad, a stop looks like high CPU and you will page yourself for an instance you turned off.
| State | When you see it in this lab |
|---|---|
| OK | The periods you asked for are not over 70, as in case 2 |
| ALARM | Both periods are over 70, as in case 1 (82 then 91) |
| INSUFFICIENT_DATA | The alarm is new, or the instance is stopped and not publishing |
Create cpu-high-lab
The alarm does not launch an instance. It only changes its own state, and it runs actions you attach. Section 17.5 attaches email. Section 17.6 attaches a scaling policy. Create the alarm now so those sections edit this alarm, not a second one with a similar name.
- CloudWatch, Alarms, Create alarm. Names may vary.
- Select metric, EC2, Per-Instance Metrics. Find the instance that serves example.com (instance A, or the one instance the group is running in section 17.6). Select CPUUtilization. Check the instance id. The wrong instance makes a perfect alarm on a computer nobody is using.
- Statistic Average, period 5 minutes. Threshold type Static. Condition Greater than, value 70. Additional configuration: 2 datapoints out of 2 evaluation periods. Missing data: Treat missing data as missing.
- Skip the notification if the SNS topic does not exist yet. You will add it in section 17.5. Alarm name:
cpu-high-lab. - Create. The state is INSUFFICIENT_DATA until two periods exist. Leave it. Do not delete it and recreate it because the state looks unfamiliar.
Lab
Chala, follow the five create steps above, then write the two cases. Do not run yes yet.
- Open the alarm
cpu-high-lab. - You should see state INSUFFICIENT_DATA or OK, not an error.
- In your notebook, write 82 for 10:00–10:05.
- Write 91 for 10:05–10:10.
- Write the result: ALARM at the end of the second period.
- Write the other case: 95, then 40.
- Write the result: stays OK.
- Leave
yesfor section 17.6, so the email and the policy fire once, together.