API is down
, resolved after 24 min. Affects API.
- Resolved. The rollback is out and the checks pass.
- Identified. A config change routed API traffic to a drained pool. Rolling back.
- Investigating. API stopped responding to checks (HTTP 502). Opened automatically.
Post-mortem
Summary
A config change at 11:18 sent API traffic to a pool that had been drained for maintenance. Requests failed with 502 for 26 minutes.
Detection
The API monitor failed at 11:19 and again at 11:20, and opened this incident.
What we are changing
Pool changes go through review. The deploy checks that a pool has healthy hosts before routing to it. The runbook for draining a pool names the routes that use it.