Runs
Executions of a backup, read live from the orchestration layer.
View as MarkdownA run is one execution of a backup. It either produces an artifact or explains why it did not.
sctl run list --backup $BACKUP_ID
sctl run get <run-id>Where runs come from
| Origin | Run ID | Notes |
|---|---|---|
| The schedule | scheduled-<backup-id>-<unix-timestamp> | Fired by the orchestration layer, not by our API |
sctl backup trigger | Same shape | An ordinary run. It does not shift the schedule |
sctl backup submit | manual-<uuidv7> | You supply the bytes |
Manual run IDs are time-ordered, so they sort chronologically by ID alone.
The ID is stable for the life of the run, including across retries. A run that retried is still the same run, with the same ID, and its artifact is filed under the ID you saw.
Phases and steps
A run has up to two phases. Each phase has steps, and each step reports its own status, attempt count and duration.
| Phase | Runs on | Steps |
|---|---|---|
produce | Your worker, or ours | The dump, and for local backups the encrypt, checksum and upload |
post-backup | Ours, always | inspect, compress, encrypt, plan, the deliveries, finalize |
A manual submission has no produce phase in the usual sense: the run waits for your upload,
then post-backup takes over.
sctl run get scheduled-018f3c2a-1754640000The step list is where a failure is actually diagnosed. "The run failed" is not useful;
"DumpPostgresActivity, attempt 2, FATAL: password authentication failed" is.
Statuses
| Status | Meaning | Artifact |
|---|---|---|
Running | In flight | Not yet |
Completed | Every step succeeded, every requested copy stored | Yes |
Failed | A step exhausted its retries, or a delivery did not land | Usually not. See below |
TimedOut | A step exceeded its timeout | No |
Terminated / Canceled | Stopped from outside | No |
The exception worth knowing
A run that stored some copies but not all is recorded as failed, with the artifact written. The artifact's locations say exactly which copies exist.
This is deliberate. A destination you selected did not get your data, so reporting success would be a claim we cannot back. But the copies that did land are real, and hiding them behind a failed run would be worse.
A run that stored nothing fails before the artifact record is written, so there is no artifact at all and a retry still has the bytes to work from.
Reading a failed run
Work down this list.
1. Which phase failed? produce is the source, the credential, or the machine. post-backup is our pipeline or your delivery targets.
2. Which step, and on which attempt? A step that failed on attempt 1 of 5 and then succeeded is not a failure at all. A step that burned every attempt with the same message is a configuration problem, not a transient one.
3. What does the failure text say? For the dump steps it is the tool's own stderr, truncated. pg_dump's message about authentication is worth more than any wrapper we could put around it.
| Failure looks like | Usually is |
|---|---|
ConfigDrift, SourceTypeMismatch | The worker's config and the definition disagree |
MissingTool | pg_dump or similar not installed on the worker |
InvalidSource | A required field missing, or a path that is not what it claims |
| Authentication errors from the tool | The source credential |
stored N of M copies | A delivery destination, not the backup |
no copy of run … was stored | Every destination failed. Check them all |
Non-retryable failures fail immediately rather than burning attempts, because retrying a misconfiguration for twenty minutes helps nobody.
See worker troubleshooting for the local side in detail.
Missing runs
A run that never appears is a different problem from one that failed, and the evidence is in a different place.
| Failed | Missing | |
|---|---|---|
| Record exists | Yes, with an error | No |
| Look at | The step list, and the worker log | The schedule, the backup state, the worker assignment |
| Usual cause | The source or the pipeline | The backup is paused or draft, or its worker is offline |
A local backup whose worker is offline produces neither. Its runs queue on that worker's queue and execute when it returns, late rather than lost.
The history window
Run history is retained for 30 days. It is read live from the orchestration layer rather than mirrored into a database, which is what makes it cheap to keep detailed step-level records, and what bounds how long they last.
Two consequences:
- Artifacts outlive their runs. An artifact under a 90-day retention still exists long after the run that produced it stopped being listed. The artifact record carries what you need: size, checksum, key ID, locations and creation time.
- An empty listing means "no history", not "nothing ran". The listing says so rather than implying the backup never worked.
If you need run outcomes retained beyond the window, send them somewhere you control as they happen. See Webhooks.
Watching a run
sctl backup trigger $BACKUP_ID
sctl run list --backup $BACKUP_ID
sctl run get <run-id>Long steps report progress rather than going silent: a dump heartbeats with the byte count so far, an upload with the part number. A step that is genuinely stuck stops heartbeating and is retried, rather than hanging until its timeout.