Backups: the 3-2-1 rule in practice
Everyone knows three copies, two media, one off-site. The rule is not wrong; it is just easy to satisfy on paper while failing in the only moment that matters. The gaps are almost always the same four.
1. Untested restores
A backup you have never restored is a hypothesis. Restoring one file is not enough — restore a whole service onto a scratch host at least once, and time it. That number is your real recovery time, not the number you wrote in the runbook.
2. Shared credentials
If the backup target is reachable with the same credentials as the source, one bad day takes both. Keep the backup destination writable by a separate identity, and keep the delete permission away from the machines being backed up.
3. Silent failure
Jobs fail quietly more often than loudly. Alert on the absence of a successful backup, not on the presence of an error message — a job that stopped running produces no error at all.
# alert on staleness, not on failure
if [ "$(find /backups -name '*.complete' -mtime -1 | wc -l)" -eq 0 ]; then
notify "no fresh backup"
fi
4. No idea what is covered
Write down what is excluded and why. The uncovered 5% is what you will discover during the incident, and the discovery always costs more than the backup would have.
The drill
Once a quarter, pick a service at random, restore it from the off-site copy with a stopwatch running, and fix whatever surprised you. The drill is the product; the backup is just the input.