---
title: "What is backup orchestration?"
description: "Any engineer can write the dump command in twenty minutes. What breaks is the scheduling, the retries, the retention, the credential handling and the restore. That gap has a name, and it is a category rather than a feature."
url: "https://saved.sh/blog/what-is-backup-orchestration"
date: "2026-08-08"
author: "saved.sh"
tag: "Engineering"
---

Writing the command that backs up a database takes about twenty minutes. `pg_dump`,
a pipe, an object storage URL. Nobody gets stuck there.

What people get stuck on is everything after the command runs for the four
hundredth time on a Tuesday, on a box nobody has logged into since March, with a
credential that expired in February, writing to a bucket whose lifecycle rule
someone changed in January.

Backup orchestration is the name for that second problem.

## The command is the smallest part [#the-command-is-the-smallest-part]

Here is the same job described two ways. The first is what people think they are
building. The second is what they end up maintaining.

| What you write                             | What you end up owning                                                               |
| ------------------------------------------ | ------------------------------------------------------------------------------------ |
| One dump command                           | Fifteen of them, one per source, each with its own flags                             |
| A cron line                                | Schedules, overlap rules, timezones, and what happens when a run outlasts its window |
| `\|` into `aws s3 cp`                      | Retries, resumption, partial uploads, and detecting a truncated dump                 |
| A bucket name                              | Retention, immutability windows, and who is allowed to delete                        |
| `AWS_SECRET_ACCESS_KEY` in the environment | Credential custody, rotation, and blast radius                                       |
| Nothing                                    | Proof that any of it worked, and a tested path back                                  |

The first column is an afternoon. The second column is a system, and it is the
same system at every company that has ever tried to build it in-house.

<Figure caption="The dump is one step. Orchestration is the machinery around it that makes the step reliable, repeatable and provable.">
  <ArtifactPipeline />
</Figure>

## A working definition [#a-working-definition]

**Backup orchestration is everything between "we should back this up" and "we
restored it".** Concretely, five jobs:

1. **Schedule.** Decide when a run happens, and what to do when the last one has
   not finished.
2. **Execute durably.** Survive a worker dying mid-run without silently skipping
   a night.
3. **Deliver.** Get the bytes to one or more destinations, compressed and
   encrypted, without staging the whole artifact on a disk that might be too small.
4. **Retain.** Decide what to keep, what to expire, and what nobody is allowed to
   delete, including us.
5. **Prove.** Leave a record that says what ran, what it produced, how large it
   was, and where the copies went.

A tool that does one of these is a feature. A tool that does all five is the
category.

## Why storage vendors are not this [#why-storage-vendors-are-not-this]

The most common confusion is between orchestration and storage, because both
end with bytes in a bucket.

Storage answers "where do the bytes live and what do they cost". Orchestration
answers "did the right bytes get made, on time, from the right source, and can
you get them back". Those are different products with different failure modes.
A storage vendor with perfect durability will happily store four hundred
consecutive copies of a zero-byte file, and its dashboard will be green the
entire time.

<Stats>
  <Stat value="5" label="Jobs orchestration has to do" />

  <Stat value="1" label="Of them a storage vendor covers" />

  <Stat value="0" unit="bytes" label="A green dashboard can be hiding" />
</Stats>

## Why we built it to hand the bytes to you [#why-we-built-it-to-hand-the-bytes-to-you]

If orchestration and storage are genuinely separate problems, then a company
solving the first has no particular reason to hold your data.

So we do not require it. Backups land in a bucket you own, under credentials you
control, encrypted before they leave the machine that produced them. We keep the
metadata that makes the five jobs possible: the schedule, the run record, the
retention policy, the list of where each copy went.

That has a consequence we are happy to be held to. Because the artifacts are
yours and the encryption happens on your side, getting your data back does not
require us to be reachable, solvent, or willing. The documented restore path
uses standard tools on files sitting in your own bucket.

<Aside title="The question worth asking any backup vendor">
  Not "how durable is your storage". Ask instead: if your company disappeared
  tonight, what exactly would I run tomorrow morning to get my data back? A vendor
  that has thought about orchestration as a separate concern from storage will
  have a real answer. One that has conflated them will talk about their SLA.
</Aside>

## Where to go from here [#where-to-go-from-here]

If you are currently on the afternoon-sized version of this problem, the honest
starting point is not to buy anything. It is to find out whether the backups you
already have can be restored, because that single test tends to reorder every
other priority.

Once you know the answer, the five jobs above are the checklist for whatever you
build or buy next.
