Writing the command that backs up a database takes about twenty minutes. pg_dump,
a pipe, an object storage URL. Nobody gets stuck there.
What people get stuck on is everything after the command runs for the four hundredth time on a Tuesday, on a box nobody has logged into since March, with a credential that expired in February, writing to a bucket whose lifecycle rule someone changed in January.
Backup orchestration is the name for that second problem.
The command is the smallest part
Here is the same job described two ways. The first is what people think they are building. The second is what they end up maintaining.
| What you write | What you end up owning |
|---|---|
| One dump command | Fifteen of them, one per source, each with its own flags |
| A cron line | Schedules, overlap rules, timezones, and what happens when a run outlasts its window |
| into aws s3 cp | Retries, resumption, partial uploads, and detecting a truncated dump |
| A bucket name | Retention, immutability windows, and who is allowed to delete |
AWS_SECRET_ACCESS_KEY in the environment | Credential custody, rotation, and blast radius |
| Nothing | Proof that any of it worked, and a tested path back |
The first column is an afternoon. The second column is a system, and it is the same system at every company that has ever tried to build it in-house.
A working definition
Backup orchestration is everything between "we should back this up" and "we restored it". Concretely, five jobs:
- Schedule. Decide when a run happens, and what to do when the last one has not finished.
- Execute durably. Survive a worker dying mid-run without silently skipping a night.
- Deliver. Get the bytes to one or more destinations, compressed and encrypted, without staging the whole artifact on a disk that might be too small.
- Retain. Decide what to keep, what to expire, and what nobody is allowed to delete, including us.
- Prove. Leave a record that says what ran, what it produced, how large it was, and where the copies went.
A tool that does one of these is a feature. A tool that does all five is the category.
Why storage vendors are not this
The most common confusion is between orchestration and storage, because both end with bytes in a bucket.
Storage answers "where do the bytes live and what do they cost". Orchestration answers "did the right bytes get made, on time, from the right source, and can you get them back". Those are different products with different failure modes. A storage vendor with perfect durability will happily store four hundred consecutive copies of a zero-byte file, and its dashboard will be green the entire time.
5
Jobs orchestration has to do
1
Of them a storage vendor covers
0bytes
A green dashboard can be hiding
Why we built it to hand the bytes to you
If orchestration and storage are genuinely separate problems, then a company solving the first has no particular reason to hold your data.
So we do not require it. Backups land in a bucket you own, under credentials you control, encrypted before they leave the machine that produced them. We keep the metadata that makes the five jobs possible: the schedule, the run record, the retention policy, the list of where each copy went.
That has a consequence we are happy to be held to. Because the artifacts are yours and the encryption happens on your side, getting your data back does not require us to be reachable, solvent, or willing. The documented restore path uses standard tools on files sitting in your own bucket.
Where to go from here
If you are currently on the afternoon-sized version of this problem, the honest starting point is not to buy anything. It is to find out whether the backups you already have can be restored, because that single test tends to reorder every other priority.
Once you know the answer, the five jobs above are the checklist for whatever you build or buy next.