Rehearsal
The drill that turns a backup into a recovery plan.
View as MarkdownA backup you have never restored is a hypothesis. Rehearsal is how it becomes a fact.
This page is a drill you can run in under an hour, and a short list of the failures it catches. Every one of them is silent until the day it is not.
The drill
Do it on a normal day, on a laptop, against a scratch target. Log the CLI out first, so you are testing recovery rather than testing your session.
1. Prove you can get the bytes.
sctl artifact list --backup $BACKUP_ID
sctl artifact get <artifact-id>
sctl artifact download <artifact-id> --output ./drill/artifactNote the Key field. That is the fingerprint you need to match against a private key you
actually hold.
2. Prove you hold the key, on a machine that is not the usual one.
gpg --list-secret-keys
gpg --list-packets ./drill/artifact | head -3The key ID in the second command's output must appear in the first. This is the step people skip and the failure that ends companies.
3. Decrypt, from nothing but the file and the key.
# Linux and macOS
gpg --decrypt ./drill/artifact > ./drill/restored
file ./drill/restored# Windows PowerShell
gpg --output .\drill\restored --decrypt .\drill\artifact
Format-Hex .\drill\restored -Count 16Run at least one rehearsal on a different operating system from the one that made the backup, if your team uses more than one. It is the cheapest way to find out that your runbook assumes tools only one person has.
4. Restore into scratch.
createdb drill_scratch
pg_restore --dbname drill_scratch --no-owner --no-privileges ./drill/restored5. Check the data, not the exit code.
psql -d drill_scratch -c '\dt'
psql -d drill_scratch -c 'SELECT count(*) FROM users;'
psql -d drill_scratch -c 'SELECT max(created_at) FROM orders;'That last query is the real output of the drill. It is your actual recovery point, and it is usually earlier than people assume, because a dump reflects the moment it started rather than the moment the run finished.
6. Tear it down.
dropdb drill_scratch
rm -rf ./drillThe failures this catches
Each of these has a plausible-looking backup sitting behind it.
| Failure | How it looks until the drill | Found at step |
|---|---|---|
| Nobody kept the private key | Runs succeed. Artifacts accumulate. All unreadable | 2 |
| The key is on one laptop, which is the one that failed | Fine, until it is not | 2 |
| The backup was never encrypted at all | Runs succeed | 1, Encrypted: no |
| Schema drift: the dump predates three migrations | The restore works and the app does not | 5 |
| Roles do not exist in the target cluster | role "app" does not exist | 4 |
pg_restore older than the dump's server | Syntax errors partway through | 4 |
| The dump is of an empty or wrong database | Small artifact, successful run | 5 |
| Nobody knows which artifact corresponds to what | Confusion at 3am | 1 |
The script source wrote a log, not a backup | Successful runs, tiny artifacts | 3 |
The two in bold are the ones that turn an incident into a company-ending event. Everything else is an afternoon.
If step 2 fails, nothing else matters. We hold ciphertext and no key. There is no support request, no override and no recovery path. Find this out on a Tuesday.
How often
| Cadence | What to do |
|---|---|
| Quarterly | The full drill above, on the backups that matter |
| After any change to retention, encryption keys, source type or worker | The full drill on that backup |
| After any change to your schema tooling | Steps 4 and 5, restore plus migrate forward |
| Monthly, at least | Step 1 and 2 only: confirm artifacts exist and the key still opens one |
The monthly version takes five minutes and catches the worst two failures. If you only ever do one thing from this page, do that.
Rehearse the whole thing, not just the restore
A restore drill that starts with "download the artifact with sctl" has assumed the part most likely to be missing.
Once a year, run the drill with a harder premise: assume we are unreachable.
- Can you get the bytes? That means a destination you own, read with your own credentials and your own S3 client.
- Can you identify the file with no metadata from us? See identifying an artifact from its bytes.
- Does the person who would be doing this know where the key is?
If the answer to the first is no, that is not a rehearsal finding, it is a design finding, and the fix is a delivery destination rather than a better runbook.
Write it down
The drill's value is in the record, not in the hour. Each time, note:
| Field | Example |
|---|---|
| Date, and who ran it | 2026-08-08, Sam |
| Backup and artifact ID | prod-db, 018f3c2a-… |
| Artifact age when restored | 14 hours |
| Actual recovery point | max(created_at) was 02:04 |
| Wall-clock time, download to verified | 38 minutes |
| Anything that surprised you | pg_restore needed --no-owner |
| Follow-up actions | Document the roles requirement in the runbook |
The two rows that get quoted back at you later are the recovery point and the wall-clock time. Those are your real RPO and RTO, as measured rather than as hoped, and they are the numbers to put in front of anyone asking what happens if the database is lost.
The first drill is always the slow one, and the surprises in it are the point. If your first rehearsal takes three hours and turns up four problems, it did its job.
Make it a calendar event
An intention to rehearse is not a rehearsal. Put it in the calendar, with a named owner, on a fixed cadence, with the checklist above attached.
The failure this guards against is not technical. It is that everyone agrees rehearsal matters, and nobody's name is on it.