What stays local
Exactly which bytes leave your network, and which never do.
View as MarkdownA local backup is worth running only if you know precisely where the boundary is. This page draws it, in both directions, including the parts that are less flattering than the summary.
The short version
| Leaves your network | Stays | |
|---|---|---|
| Source credentials | No | config.yaml, on your disk |
| Private encryption key | No | We never had it |
| Plaintext dump | No | A scratch file, deleted after upload |
| Encrypted artifact | Yes, to object storage | |
| Checksum, byte count, filename | Yes, to our API | |
| Identifiers, progress, error text | Yes, to run history |
What never leaves
Source credentials. Database passwords, API tokens, bucket keys and script paths live in
the worker's config.yaml and are read by the worker process. They are never sent to us, and
there is no code path that could: a run task carries three identifiers, and the worker
resolves everything else from its own file.
The private half of your encryption key. We hold the public half only, and only because you put it in the config. This is not a policy, it is an absence of capability. We cannot decrypt your backups, and we cannot recover them if you lose the key.
The plaintext dump. It exists as a file under local_temp_path between the dump step and
the encrypt step, on your machine, in a directory created 0700. It is removed when the run
ends, including when the run fails.
The worker key itself. We store its identifier so we know which worker is calling. The key is issued by the identity provider and held by you. It is not in our database and not in our logs.
Passwords are passed to pg_dump and mysqldump through the environment,
never on the command line, so they do not appear in ps output to other users on the host.
The config file's permissions are still the thing protecting them. Keep it 0600.
What we receive
Three channels, and it is worth being precise about each.
The artifact
Uploaded directly from your machine to object storage using a presigned URL scoped to that one run. It does not pass through our API, and the worker never holds long-lived storage credentials.
It is encrypted only if that backup has a key. encryption.public_key is per backup and optional, and
without it the worker uploads what came out of the dump. There is no server-side encryption
standing behind that, because the design deliberately has no key to do it with. See
encryption.
The confirmation
One API call at the end of a run, carrying:
| Field | Example |
|---|---|
| Run ID | scheduled-<backup-id>-<timestamp> |
| Checksum | sha256:9f86d0…, taken after encryption |
| Size | Byte count of the uploaded file |
| Filename | app.dump.gz.gpg |
The filename is derived from the source: the database name, the bucket name, the folder or file's base name, the script's own name, or the last path segment of a URL. Each layer the run applied adds its own suffix, so the name says how to unwrap it. Names are not secret to us. If a database name is itself sensitive, that is worth knowing before you schedule it.
Run history
The orchestration layer carries control information only, but "only control information" is still more than nothing:
| What | Detail |
|---|---|
| Identifiers | Workspace, backup and run IDs |
| Scratch paths | Progress messages name the local staging file being worked on |
| Progress | Byte counts, object counts, upload part numbers |
| Error text | Up to 2000 characters of the failing tool's stderr |
That last row is the one to think about. When pg_dump fails, its own message is what we keep,
because "access denied for user backup@db.internal" is worth vastly more to you than "exit
status 2". It can contain hostnames, usernames, database and table names. It does not contain
the password, which never reaches the command line, and it is truncated so one broken backup
cannot flood the history.
If that trade is wrong for you, the lever is the script source: run your own command, handle
its errors yourself, and write only what you want to $SAVED_OUTPUT. What the worker reports
is then whatever your script chose to print.
The plaintext window
Between the dump completing and the encryption completing, an unencrypted copy of your data exists on the worker host. This is unavoidable: something has to produce the bytes before something else can encrypt them.
What bounds it:
- Location.
local_temp_path, default<os temp>/saved, created0700and owned by the worker's user. - Lifetime. Removed as soon as the run finishes, in a cleanup step that runs even when the run failed or was cancelled.
- Your choice of host. The worker should run somewhere you would already be comfortable holding a database dump, because for a few minutes per run it does.
Set local_temp_path to a filesystem you control and can size. Encrypted volumes are a
reasonable choice here, and cost you nothing but throughput.
Keeping even the ciphertext home
If storing ciphertext with us is more than you want, a backup can deliver to your own bucket instead of, or in addition to, our storage. The worker's behaviour is unchanged; what changes is where the artifact lands.
See Delivery.
Comparing against a cloud backup
The same source type can run either way, and the difference is entirely about who holds the credential.
| Local | Cloud | |
|---|---|---|
| Who connects to the source | Your worker | We do |
| Where the source credential lives | Your config.yaml | Our vault, read at run time |
| Where plaintext exists | Your host | Our infrastructure |
| Where encryption happens | Your host | Our infrastructure |
| Needs a running process | Yes | No |
Local is the stronger position and the one to prefer when you can run a worker. Cloud exists because sometimes you would rather we connected than stand up a process, and that is a legitimate trade as long as it is made explicitly.
script, file and folder are local-only, permanently. We do not execute customer scripts
in our cloud, and your filesystem does not exist there.