Pick a target and a clock
Choose one production database and a scratch server. Start a timer. The goal is not perfection — it is discovering where the process stalls while nothing is actually on fire.
Restore, verify, record
Restore the most recent snapshot, then verify: row counts on the five largest tables, the latest timestamp in your busiest table, and one application-level smoke query. Record how long each step took and who had to be involved.
Anything that required tribal knowledge goes into the runbook. Anything that took longer than your RTO becomes next quarter's engineering work.
Repeat quarterly
Drills decay. Schemas change, servers get replaced, people leave. A quarterly cadence keeps the runbook honest and keeps recovery boring, which is exactly what you want it to be.
