Users said the site wouldn't load. Browser: certificate expired. Last successful renew? Three months ago. For a personal site maybe forgivable — for production, no.

Cause: the cron job was left on the old server after migrating to a new one. Nobody checked because "it used to work". Automation without monitoring is just a broken cron that never rings.

Nginx was up with no obvious errors. From a process view everything looked healthy — the problem was at the TLS layer, not HTTP. openssl s_client was the first tool that showed the truth.

Manual renew that day worked, but we had ~20 minutes downtime because the certificate path in Nginx config didn't match certbot's path. A wrong symlink can extend an incident.

After the incident: certbot with a systemd timer instead of cron — easier to inspect with systemctl list-timers and clearer logs.

We set an alert 14 days before expiry. Sounds early until you've had a long holiday and a busy deploy week — then it's barely enough.

We wrote a one-page runbook: check expiry, manual renew, reload Nginx, verify with curl and openssl. Simple, but when you're panicking, defined steps help.

Main lesson: infrastructure without monitoring is gambling. SSL expiry is 100% preventable — you just have to monitor the things that are "automatic" too.