Rollback after a deploy that broke everything

Merge to main, green pipeline, successful deploy — 10 minutes later error rate went from 0.1% to 15%. Users got logged out, sessions broke, and support was flooding the group chat.
Root cause: a migration change never tested on staging — staging had an older database and the migration only ran cleanly against main's fresher schema. Pipeline was green; real data was broken.
Rollback was manual: find the previous image in the registry, retag, redeploy. 45 minutes of pure stress. No "revert to last version" button — just kubectl and hope.
Worst part: for 45 minutes we didn't know if the migration was reversible. Backend and DevOps read logs simultaneously and guessed together. Incidents without runbooks move slowly.
After the incident we changed structure: immutable image tags (never latest in production), two deploy slots (simple blue-green), and mandatory staging smoke tests with realistic data.
We added feature flags for high-risk changes — not everything, just where rollback is hard: auth, payments, migrations. Low setup cost; value shows on the first incident.
We wrote a one-page runbook: who is incident commander, how to rollback, where to read logs, when to notify users. Simple, but at 11 PM when everyone's tense, brains need checklists.
A green pipeline isn't enough — you must be healthy after deploy. Every release now includes a 15-minute monitoring window where the deployer watches error rate and latency. Anomaly means rollback without a meeting.