Configuration
One file: the worker credential, and the sources this worker serves.
View as Markdownconfig.yaml is the only configuration input, and it has exactly two required keys.
api_url: https://api.saved.sh
token: "<shown once when the worker was provisioned>"Everything else, the hub address, the namespace, the task queue, is fetched from the API at startup with that same token. None of it is a value you could know in advance, so none of it is asked for.
Where the file lives
The worker reads config.yaml relative to its working directory. There is no flag and no
path argument.
| How you run it | Where the file goes |
|---|---|
| Directly | The directory you run it from |
| systemd | Whatever WorkingDirectory points at |
| launchd | Whatever the plist's WorkingDirectory key points at |
| Task Scheduler | The action's working directory |
| Container | /config.yaml, because the image sets no working directory |
Every top-level key can also be supplied as an environment variable of the same name in
upper case: API_URL, TOKEN, DEBUG, LOCAL_TEMP_PATH, ENCRYPTION_PUBLIC_KEY,
HUB_ADDRESS and so on. When those are set the file is optional, which is how the container
image is configured.
The backups: map is the exception: it cannot come from the environment. A worker with no
file has no source configuration, and every scheduled run against it fails with
ConfigDrift.
Permissions
This file holds your database passwords. Nothing else on the host should be able to read it.
sudo install -d -m 0750 -o saved -g saved /etc/saved
sudo install -m 0600 -o saved -g saved config.yaml /etc/saved/config.yamlThe token never enters our database, and never appears in a log line. It is on your disk and nowhere else, which means the file permissions are the whole of its protection.
Full reference
Only api_url and token are required. Everything else is shown with the value the worker
uses when you leave it out.
api_url: https://api.saved.sh
token: "<shown once when the worker was provisioned>"
debug: false # verbose logging, to stderr
local_temp_path: /tmp/saved # staging dir, default <os temp>/saved
# Absolute paths to the dump tools. Only needed when they are not on PATH.
tools:
pg_dump: /usr/local/bin/pg_dump
mysqldump: /usr/local/bin/mysqldump
curl: /usr/bin/curl
# One entry per backup, keyed by the backup id from the dashboard.
backups:
"<backup-uuid>":
source_type: postgres
compression: true # gzip before upload
encryption: # per backup, not per worker
public_key: ./keys/prod.asc # a path, or the armoured key inline
source:
database: app
host: db.internal
port: 5432
user: app
password: "<password>"
ssl_mode: requireThe orchestration endpoint is not configured here. The worker fetches it from the API at startup, along with its namespace and task queue.
| Key | Default | What it does |
|---|---|---|
api_url | required | Our API. Trailing slashes are trimmed |
token | required | The worker key from provisioning |
debug | false | Debug-level logging instead of info |
local_temp_path | <os temp>/saved | Where dumps and encrypted copies are staged. Created 0700 |
encryption_public_key | none | Worker-wide fallback key, armoured block only |
tools.* | look up on PATH | Absolute path to a dump tool |
backups.<id>.source_type | required | Which source this backup is |
backups.<id>.source | required | The source's own fields, flat |
backups.<id>.compression | false | gzip the dump before upload |
backups.<id>.encryption.public_key | none | The armoured PGP public key, or a path to it. Without it, this backup uploads unencrypted |
Five more keys exist and are overrides you should not normally set, because the worker is
told these at startup. They are flat, not nested under a hub: block.
| Key | Default | What it does |
|---|---|---|
hub_tls_enabled | true | Turning this off sends the token over a cleartext connection |
hub_tls_ca_path | none, meaning the system trust store | A CA bundle, for a TLS-inspecting proxy |
hub_tls_server_name | from the address | SNI override |
hub_address | from the API | Point the worker at a different hub |
hub_namespace | from the API | Point the worker at a different namespace |
Encryption
encryption is optional, and leaving it out means that backup's artifacts are uploaded as
they came out of the dump. The worker does not refuse to run without a key; it just skips
the encrypt step. Set it on every backup you care about.
Encryption is per backup, not per worker. One machine can hold backups for different systems, and the key that protects one is not the key for another.
public_key takes either the armoured public half inline or a path to a file holding it.
A relative path resolves against the config file's own directory, so the config is runnable
from anywhere. A path that cannot be read, or a file that holds no armoured key, fails at
startup: skipping it would upload the artifact in the clear, which is the one failure you
would not notice.
The worker-wide fallback
A top-level encryption_public_key is used by any backup that sets no key of its own. It is
a convenience for a worker whose backups all belong to the same system, and the per-backup
key always wins where both are set.
The fallback takes the armoured block only. Unlike the per-backup key, a path is not
resolved and not checked at startup. Give it a file path and every run fails at the encrypt
step with read public key, rather than at startup where you would have seen it.
encryption_public_key: |
-----BEGIN PGP PUBLIC KEY BLOCK-----
...
-----END PGP PUBLIC KEY BLOCK-----Generate one, keep the private half somewhere you will still have it after the incident that makes you need it, and give the worker the public half:
gpg --quick-generate-key "backups@example.com" default default never
gpg --armor --export backups@example.comWe hold only the public key, and only because you put it here. We cannot decrypt your
backups, and we cannot recover them if you lose the private key. Encryption happens on this
machine, before anything is uploaded. When compression is on, the dump is gzipped first, so
the artifact is compressed whether or not it is also encrypted.
Connecting to the hub
The hub is not configured here. At startup the worker calls the API with its token and is
told its address, namespace, worker ID and task queue. A worker with a valid api_url and
token therefore needs nothing else to connect, and moving the hub does not mean editing
every worker's file.
TLS is on by default and is what hub.saved.sh expects. The system trust store is used, so
there is no CA file to supply: hub.saved.sh presents a publicly trusted certificate.
External tools
Four of the source types shell out to a tool that has to be on the host. The s3, file,
folder and script sources need nothing beyond the binary itself.
source_type | Tool | tools key |
|---|---|---|
postgres | pg_dump | pg_dump |
mysql | mysqldump | mysqldump |
redis | none | |
web | curl | curl |
The worker resolves all four at startup and logs the result. A missing tool is a warning,
not a fatal error: it fails only the runs of the source type that needs it, with a
non-retryable MissingTool. That is deliberate, so a host that only does file backups does
not need a Postgres client.
The container image ships all four.
Per-backup sources
Every backup this worker serves needs an entry under backups:, keyed by the backup ID
from the dashboard or sctl backup list.
backups:
"018f3c2a-9e11-7c4d-b0a1-2e6f5d3c9a70":
source_type: postgres
source:
database: app
host: db.internal
port: 5432Three failure modes are non-retryable on purpose, because the worker will not guess:
| Error | Cause |
|---|---|
ConfigDrift | A run fired for a backup with no entry here. Add it, or remove the schedule |
SourceTypeMismatch | The entry's source_type disagrees with the workflow that fired |
InvalidSource | The source: block is missing a required field, or holds a key that source type does not have |
Write the natural YAML type. A port is a number, a flag is a boolean, and a list is a
list. The worker parses source: into a struct per source type, so it knows what each field
should be. Quoted forms are still accepted, so an older config keeps working.
port: 5432 # not "5432"
no_owner: true # not "true"
exclude_tables: [audit_log, sessions]An unknown key under source: fails at startup, rather than being ignored. A misspelled
field would otherwise look applied and quietly do nothing, which you would discover during a
restore. The error names the key and the source type.
exclude, include, args, tables and exclude_tables are lists. A newline-separated
block scalar is still read as one entry per line, so configs written before these were lists
parse unchanged.
credentials: was renamed to source:, and the fields inside it did not change. A
config still using the old name fails at startup with a message saying so, rather than
starting with no source configuration.
Source types
Local backups support eight source types. script, file and folder are local-only,
permanently: they run your commands or read your own disk, neither of which our cloud can or
should do.
postgres
Runs pg_dump -Fc (custom format), producing <database>.dump.
| Key | Required | Notes |
|---|---|---|
database | yes | |
host, port, user | no | Omitted flags fall back to pg_dump's own defaults |
password | no | Passed as PGPASSWORD, never on the command line |
ssl_mode | no | Passed as PGSSLMODE, e.g. require, verify-full |
schema | no | Restrict to one schema |
exclude_tables | no | A list of table names |
no_owner | no | true adds --no-owner |
no_privileges | no | true adds --no-privileges |
"<backup-uuid>":
source_type: postgres
source:
database: app
host: db.internal
port: 5432
user: backup
password: "<password>"
ssl_mode: require
exclude_tables: [audit_log, sessions]
no_owner: true
no_privileges: trueThe worker always passes --no-password, so pg_dump fails fast with a readable error
instead of blocking on an interactive prompt no one will ever answer.
mysql
Runs mysqldump --single-transaction --quick --routines --events --triggers, producing
<database>.sql.
| Key | Required | Notes |
|---|---|---|
database | yes | |
host, port, user | no | |
password | no | Passed as MYSQL_PWD, never on the command line |
ssl_mode | no | Passed through as --ssl-mode |
tables | no | A list. Restrict to specific tables |
exclude_tables | no | A list. Becomes --ignore-table=<db>.<table> |
no_data | no | true dumps schema only |
Setting both tables and exclude_tables is refused at startup. They express opposite
intentions and the worker will not pick one.
--single-transaction takes a consistent snapshot without locking the whole database, which
is the difference between a nightly backup you can run in production and one you cannot.
redis
Reads every key with SCAN and DUMP, producing <host>.redis.jsonl. Needs no external tool, and works against managed Redis.
| Key | Required | Notes |
|---|---|---|
host, port | no | Also names the artifact; defaults to redis.redis.jsonl |
user, password | no | Password passed as REDISCLI_AUTH |
db | no | Database index |
tls | no | true adds --tls |
insecure | no | true skips certificate verification. Refused unless tls is also set |
A server that refuses the sync exits zero and writes nothing. The worker treats an empty snapshot as a failure rather than uploading a zero-byte artifact.
s3
Copies a bucket or prefix into one tar, producing <bucket>.tar. Any S3-compatible store
works: AWS S3, Cloudflare R2, MinIO, Backblaze B2, Wasabi.
| Key | Required | Notes |
|---|---|---|
bucket | yes | |
access_key_id | yes | |
secret_access_key | yes | |
region | no | Defaults to auto, which is what R2 and several others want |
endpoint | no | For non-AWS stores. https:// is added if you omit the scheme |
prefix | no | Restrict to one path |
use_path_style | no | true for MinIO and most self-hosted stores |
"<backup-uuid>":
source_type: s3
source:
bucket: uploads
endpoint: minio.internal:9000
access_key_id: "<key>"
secret_access_key: "<secret>"
use_path_style: trueAn object rewritten while it is being copied fails the run. The tar header is written from
the listing and cannot be corrected afterwards, so the alternative would be a silently
truncated file, which is worse than a failure you can see. An empty bucket or prefix fails
with EmptySource rather than producing an empty archive.
web
Fetches a URL with curl and stores the response body.
| Key | Required | Notes |
|---|---|---|
url | yes | Must start with https://. An http:// URL is refused rather than silently upgraded |
method | no | Defaults to curl's own, so GET |
headers | no | One Name: value per line |
"<backup-uuid>":
source_type: web
source:
url: https://api.internal/export
headers: |
Authorization: Bearer <token>
Accept: application/jsonRedirects are followed, but only to http and https, so a redirect to file:// cannot
turn a web backup into a filesystem read. HTTP error statuses fail the run. The artifact is
named from the response's Content-Disposition, falling back to the last path segment of the
final URL.
Values must not contain a line break. The worker rejects one rather than risk a newline becoming a second directive.
script
Runs a command you own and stores whatever it writes to the path in $SAVED_OUTPUT.
| Key | Required | Notes |
|---|---|---|
path | yes | Must be executable, unless interpreter is set |
args | no | A list, one entry per argument |
interpreter | no | Runs <interpreter> <path> <args...> |
working_dir | no | Directory to run in |
"<backup-uuid>":
source_type: script
source:
path: /opt/saved/dump-everything.sh
working_dir: /opt/saved#!/usr/bin/env bash
set -euo pipefail
pg_dump --format=custom mydb > "$SAVED_OUTPUT"Write to $SAVED_OUTPUT, not to stdout. Anything on stdout or stderr is treated as a
log and kept for diagnostics. A script that exits zero without writing to that path fails
with EmptyOutput, which is a better outcome than an empty artifact you discover during a
restore.
file
Copies one regular file, keeping its name.
| Key | Required | Notes |
|---|---|---|
path | yes | |
follow_symlinks | no | true to back up a symlink's target |
A symlink without follow_symlinks is refused rather than quietly backing up the link text.
Directories, sockets, devices and FIFOs are refused with an explanation.
folder
Walks a directory tree into a zip, producing <folder>.zip.
| Key | Required | Notes |
|---|---|---|
path | yes | |
include | no | A list of globs. Empty means everything |
exclude | no | A list of globs. Applied before include |
follow_symlinks | no | true resolves links; otherwise they are stored as links |
"<backup-uuid>":
source_type: folder
source:
path: /srv/uploads
exclude: ["*.tmp", "node_modules", ".git"]A pattern matches against the full relative path, the base name, or any single path segment,
so node_modules excludes the directory wherever it appears. Excluding a directory skips the
whole subtree. Entries are stored without zip compression, because the encryption step
compresses anyway and doing it twice costs CPU for nothing.
When follow_symlinks is set, each resolved target is archived once, so a link cycle cannot
produce an infinite archive.
Multiple workers on one host
Each worker needs its own working directory, because each needs its own config.yaml.
/etc/saved/prod/config.yaml # worker A
/etc/saved/staging/config.yaml # worker BRun one process per directory. They can share a local_temp_path, since staging files are
uniquely named, but giving each its own makes disk pressure easier to attribute.
What you must not do is run the same config twice. See one process per credential.