Backups: the 3-2-1 rule in practice

Everyone knows three copies, two media, one off-site. The rule is not wrong; it is just easy to satisfy on paper while failing in the only moment that matters. The gaps are almost always the same four.

1. Untested restores

A backup you have never restored is a hypothesis. Restoring one file is not enough — restore a whole service onto a scratch host at least once, and time it. That number is your real recovery time, not the number you wrote in the runbook.

2. Shared credentials

If the backup target is reachable with the same credentials as the source, one bad day takes both. Keep the backup destination writable by a separate identity, and keep the delete permission away from the machines being backed up.

3. Silent failure

Jobs fail quietly more often than loudly. Alert on the absence of a successful backup, not on the presence of an error message — a job that stopped running produces no error at all.

# alert on staleness, not on failure
if [ "$(find /backups -name '*.complete' -mtime -1 | wc -l)" -eq 0 ]; then
  notify "no fresh backup"
fi

4. No idea what is covered

Write down what is excluded and why. The uncovered 5% is what you will discover during the incident, and the discovery always costs more than the backup would have.

The drill

Once a quarter, pick a service at random, restore it from the off-site copy with a stopwatch running, and fix whatever surprised you. The drill is the product; the backup is just the input.

← Back to all notes