The Slack #alerts channel was spam. CPU above 70%? Alert. One queue retry? Alert. Deployment finished? Alert. When the disk actually filled up, nobody noticed — the last 50 messages were all "informational".

One day error rate genuinely spiked and it took 20 minutes for anyone to react. Not indifference — alert fatigue. The team had learned to ignore the channel.

We sat down and set one simple rule: an alert must be actionable — someone should get up at 3 AM and do a specific thing. If the answer is "wait, it'll fix itself", it's not an alert.

We moved informational signals to dashboards, not Slack. Deployment success, CPU at 65%, pod restart after deploy — all go to Grafana. Slack is only for things that need immediate action.

We added severity: P1 (users affected), P2 (degraded), P3 (look this week). P1 pages on-call, P3 opens a ticket. Without severity everything feels urgent.

We tuned thresholds against real baselines — not magic numbers from docs. Collected a week of metrics first, then set alerts. 70% CPU meant nothing for a lightweight service; it was normal for a batch job.

We set up on-call rotation — even on a small team. One person owns P1/P2 for their week. Less burnout because not everyone feared their phone vibrating at night.

Alert quality beats quantity. After cleanup, alert volume dropped ~85% — but real incidents get caught faster. Makes sense: when the channel is quiet, a new message actually gets seen.