At 2 AM: service down. Logs full of connection refused. Cause? A small docker-compose change mounted a volume to the wrong path — the database pointed at an empty folder instead of /var/lib/postgresql/data.

We had no health checks and restart policy was always — the container kept restarting without anyone seeing the root cause. From the outside the service looked "up" but requests timed out.

First move was rollback, but the latest backup snapshot was three days old. Three days of data gone — not from intentional deletion, but because the app started on an empty volume and initialized a fresh database.

After the incident we set a few simple rules: no docker-compose change without a second review, no deploy without an automated pre-merge backup, and staging must match production volume paths exactly — not "close enough".

We added healthchecks to every critical service. Postgres uses pg_isready; the API exposes /health that actually hits the DB — not a hollow 200 OK.

A technical detail people skip: depends_on only guarantees start order, not readiness. Without healthchecks, startup race conditions are basically guaranteed.

Every deploy now includes a post-start smoke test: a simple DB query, an API request, and a check of the mount point inside the container. Adds maybe 30 seconds, saves a 2 AM wake-up.

The most expensive lessons cost time and nerves. But if you walk away with a solid checklist, at least you won't repeat the same mistake.