---
title: "What a run record is for"
description: "A directory of files tells you a backup happened. A run record tells you what happened, whether it was complete, and when it stopped being true."
url: "https://saved.sh/blog/what-a-run-record-is-for"
date: "2026-07-26"
author: "saved.sh"
tag: "Engineering"
---

Most homegrown backup setups produce a directory. Files with dates in the names,
sorted newest first, and a rough sense that things are fine because the directory
is not empty.

A directory answers one question: did something get written. It cannot answer the
three that matter when you are staring at it during an incident.

<Stats>
  <Stat value="1" label="Is this file complete, or did the dump die halfway?" />

  <Stat value="2" label="Is it the size it should be, or the size of an error message?" />

  <Stat value="3" label="Which run produced it, and did that run succeed?" />
</Stats>

## The empty-file failure [#the-empty-file-failure]

This is the one that gets people, and it is worth describing precisely because
it is so undramatic.

```bash
pg_dump "$DATABASE_URL" | gzip > backup.sql.gz
```

If `pg_dump` fails, it writes an error to stderr and exits non-zero. But the
pipeline's exit status is `gzip`'s, and `gzip` succeeded: it compressed zero
bytes into a small, perfectly valid archive. The file exists. It has today's
date. It is 20 bytes.

Nothing about the directory listing looks wrong. It will not look wrong tomorrow
either, or in eight months when you need it.

<Aside title="Why this survives so long" tone="accent">
  Every individual night looks the same as a successful one. The failure has no
  signal of its own, so it is only discovered by the one action nobody performs on
  a healthy system: actually opening the file.
</Aside>

## What we record instead [#what-we-record-instead]

Every run produces a record, and the artifact is measured rather than assumed.

<Figure caption="Each stop is recorded. The size and checksum are taken at seal time, not inferred from the file later.">
  <ArtifactPipeline />
</Figure>

The record carries what the run did, how long it took, the bytes that arrived,
the checksum of what was sealed, and which worker executed it. That turns the
three unanswerable questions into a table lookup.

A backup that produced 4 KB where it produced 4.2 GB yesterday is visible in the
run history the next morning, not in a year. Not because anything clever is
inspecting the contents, but because the number is written down next to the
number from last time.

## Absence is a signal too [#absence-is-a-signal-too]

The subtler thing a run record gives you is the ability to notice a run that did
not happen.

A cron entry that stops firing produces nothing: no file, no error, no log line,
because the thing that would have written them never started. There is no
artifact of its absence. A schedule that expects a run can tell the difference
between "ran and failed" and "never ran", and those need different responses.

<Figure caption="An interrupted run under a durable engine resumes and completes. Under cron it is simply gone, and gone leaves no evidence.">
  <DurableRun />
</Figure>

## What it is not [#what-it-is-not]

A run record is not verification that the backup restores. Nothing short of
restoring it is that, and we would not claim otherwise. What it does is remove
the failures that are detectable without a restore, which is most of them, and
leave you with the one honest remaining question.

<Aside title="Still worth doing yourself">
  Restore something into a scratch environment and diff the row counts. Metadata
  catches the truncated file and the run that never fired. It cannot catch a dump
  that is complete, well-formed, and of the wrong database.
</Aside>

The directory was never lying to you. It just was not saying anything.
