---
title: "Configuration"
description: "One file: the worker credential, and the sources this worker serves."
url: "https://saved.sh/docs/workers/configuration"
---

`config.yaml` is the only configuration input, and it has exactly two required keys.

```yaml title="config.yaml"
api_url: https://api.saved.sh
token: "<shown once when the worker was provisioned>"
```

Everything else, the hub address, the namespace, the task queue, is fetched from the API at
startup with that same token. None of it is a value you could know in advance, so none of it
is asked for.

## Where the file lives [#where-the-file-lives]

The worker reads `config.yaml` **relative to its working directory**. There is no flag and no
path argument.

| How you run it                                  | Where the file goes                                         |
| ----------------------------------------------- | ----------------------------------------------------------- |
| Directly                                        | The directory you run it from                               |
| [systemd](/docs/workers/install/linux)          | Whatever `WorkingDirectory` points at                       |
| [launchd](/docs/workers/install/macos)          | Whatever the plist's `WorkingDirectory` key points at       |
| [Task Scheduler](/docs/workers/install/windows) | The action's working directory                              |
| [Container](/docs/workers/install/container)    | `/config.yaml`, because the image sets no working directory |

Every top-level key can also be supplied as an **environment variable of the same name in
upper case**: `API_URL`, `TOKEN`, `DEBUG`, `LOCAL_TEMP_PATH`, `ENCRYPTION_PUBLIC_KEY`,
`HUB_ADDRESS` and so on. When those are set the file is optional, which is how the container
image is configured.

<Callout type="warn">
  The `backups:` map is the exception: it cannot come from the environment. A worker with no
  file has no source configuration, and every scheduled run against it fails with
  `ConfigDrift`.
</Callout>

## Permissions [#permissions]

This file holds your database passwords. Nothing else on the host should be able to read it.

```bash
sudo install -d -m 0750 -o saved -g saved /etc/saved
sudo install -m 0600 -o saved -g saved config.yaml /etc/saved/config.yaml
```

The token never enters our database, and never appears in a log line. It is on your disk and
nowhere else, which means the file permissions are the whole of its protection.

## Full reference [#full-reference]

Only `api_url` and `token` are required. Everything else is shown with the value the worker
uses when you leave it out.

```yaml title="config.yaml"
api_url: https://api.saved.sh
token: "<shown once when the worker was provisioned>"

debug: false                        # verbose logging, to stderr
local_temp_path: /tmp/saved         # staging dir, default <os temp>/saved

# Absolute paths to the dump tools. Only needed when they are not on PATH.
tools:
  pg_dump: /usr/local/bin/pg_dump
  mysqldump: /usr/local/bin/mysqldump
  curl: /usr/bin/curl

# One entry per backup, keyed by the backup id from the dashboard.
backups:
  "<backup-uuid>":
    source_type: postgres
    compression: true               # gzip before upload
    encryption:                     # per backup, not per worker
      public_key: ./keys/prod.asc   # a path, or the armoured key inline
    source:
      database: app
      host: db.internal
      port: 5432
      user: app
      password: "<password>"
      ssl_mode: require
```

The orchestration endpoint is **not** configured here. The worker fetches it from the API at
startup, along with its namespace and task queue.

| Key                                  | Default           | What it does                                                                                  |
| ------------------------------------ | ----------------- | --------------------------------------------------------------------------------------------- |
| `api_url`                            | required          | Our API. Trailing slashes are trimmed                                                         |
| `token`                              | required          | The worker key from provisioning                                                              |
| `debug`                              | `false`           | Debug-level logging instead of info                                                           |
| `local_temp_path`                    | `<os temp>/saved` | Where dumps and encrypted copies are staged. Created `0700`                                   |
| `encryption_public_key`              | none              | Worker-wide fallback key, [armoured block only](#the-worker-wide-fallback)                    |
| `tools.*`                            | look up on `PATH` | Absolute path to a dump tool                                                                  |
| `backups.<id>.source_type`           | required          | Which source this backup is                                                                   |
| `backups.<id>.source`                | required          | The source's own fields, flat                                                                 |
| `backups.<id>.compression`           | `false`           | gzip the dump before upload                                                                   |
| `backups.<id>.encryption.public_key` | none              | The armoured PGP public key, or a path to it. **Without it, this backup uploads unencrypted** |

Five more keys exist and are **overrides you should not normally set**, because the worker is
told these at startup. They are flat, not nested under a `hub:` block.

| Key                   | Default                              | What it does                                                 |
| --------------------- | ------------------------------------ | ------------------------------------------------------------ |
| `hub_tls_enabled`     | `true`                               | Turning this off sends the token over a cleartext connection |
| `hub_tls_ca_path`     | none, meaning the system trust store | A CA bundle, for a TLS-inspecting proxy                      |
| `hub_tls_server_name` | from the address                     | SNI override                                                 |
| `hub_address`         | from the API                         | Point the worker at a different hub                          |
| `hub_namespace`       | from the API                         | Point the worker at a different namespace                    |

## Encryption [#encryption]

<Callout type="error">
  **`encryption` is optional, and leaving it out means that backup's artifacts are uploaded as
  they came out of the dump.** The worker does not refuse to run without a key; it just skips
  the encrypt step. Set it on every backup you care about.
</Callout>

Encryption is **per backup**, not per worker. One machine can hold backups for different
systems, and the key that protects one is not the key for another.

`public_key` takes either the armoured public half inline or a **path** to a file holding it.
A relative path resolves against the config file's own directory, so the config is runnable
from anywhere. A path that cannot be read, or a file that holds no armoured key, **fails at
startup**: skipping it would upload the artifact in the clear, which is the one failure you
would not notice.

### The worker-wide fallback [#the-worker-wide-fallback]

A top-level `encryption_public_key` is used by any backup that sets no key of its own. It is
a convenience for a worker whose backups all belong to the same system, and the per-backup
key always wins where both are set.

<Callout type="warn">
  **The fallback takes the armoured block only.** Unlike the per-backup key, a path is not
  resolved and not checked at startup. Give it a file path and every run fails at the encrypt
  step with `read public key`, rather than at startup where you would have seen it.
</Callout>

```yaml title="config.yaml"
encryption_public_key: |
  -----BEGIN PGP PUBLIC KEY BLOCK-----
  ...
  -----END PGP PUBLIC KEY BLOCK-----
```

Generate one, keep the private half somewhere you will still have it after the incident that
makes you need it, and give the worker the public half:

```bash
gpg --quick-generate-key "backups@example.com" default default never
gpg --armor --export backups@example.com
```

We hold only the public key, and only because you put it here. We cannot decrypt your
backups, and we cannot recover them if you lose the private key. Encryption happens on this
machine, before anything is uploaded. When `compression` is on, the dump is gzipped first, so
the artifact is compressed whether or not it is also encrypted.

## Connecting to the hub [#connecting-to-the-hub]

**The hub is not configured here.** At startup the worker calls the API with its token and is
told its address, namespace, worker ID and task queue. A worker with a valid `api_url` and
`token` therefore needs nothing else to connect, and moving the hub does not mean editing
every worker's file.

TLS is on by default and is what `hub.saved.sh` expects. The system trust store is used, so
there is no CA file to supply: `hub.saved.sh` presents a publicly trusted certificate.

## External tools [#external-tools]

Four of the source types shell out to a tool that has to be on the host. The `s3`, `file`,
`folder` and `script` sources need nothing beyond the binary itself.

| `source_type` | Tool        | `tools` key |
| ------------- | ----------- | ----------- |
| `postgres`    | `pg_dump`   | `pg_dump`   |
| `mysql`       | `mysqldump` | `mysqldump` |
| `redis`       | none        |             |
| `web`         | `curl`      | `curl`      |

The worker resolves all four at startup and logs the result. A missing tool is a **warning,
not a fatal error**: it fails only the runs of the source type that needs it, with a
non-retryable `MissingTool`. That is deliberate, so a host that only does file backups does
not need a Postgres client.

The container image ships all four.

## Per-backup sources [#per-backup-sources]

Every backup this worker serves needs an entry under `backups:`, keyed by the **backup ID**
from the dashboard or `sctl backup list`.

```yaml
backups:
  "018f3c2a-9e11-7c4d-b0a1-2e6f5d3c9a70":
    source_type: postgres
    source:
      database: app
      host: db.internal
      port: 5432
```

Three failure modes are non-retryable on purpose, because the worker will not guess:

| Error                | Cause                                                                                          |
| -------------------- | ---------------------------------------------------------------------------------------------- |
| `ConfigDrift`        | A run fired for a backup with no entry here. Add it, or remove the schedule                    |
| `SourceTypeMismatch` | The entry's `source_type` disagrees with the workflow that fired                               |
| `InvalidSource`      | The `source:` block is missing a required field, or holds a key that source type does not have |

**Write the natural YAML type.** A port is a number, a flag is a boolean, and a list is a
list. The worker parses `source:` into a struct per source type, so it knows what each field
should be. Quoted forms are still accepted, so an older config keeps working.

```yaml
port: 5432                        # not "5432"
no_owner: true                    # not "true"
exclude_tables: [audit_log, sessions]
```

<Callout type="warn">
  **An unknown key under `source:` fails at startup**, rather than being ignored. A misspelled
  field would otherwise look applied and quietly do nothing, which you would discover during a
  restore. The error names the key and the source type.
</Callout>

`exclude`, `include`, `args`, `tables` and `exclude_tables` are lists. A newline-separated
block scalar is still read as one entry per line, so configs written before these were lists
parse unchanged.

<Callout>
  `credentials:&#x60; was renamed to &#x2A;*`source:`**, and the fields inside it did not change. A
  config still using the old name fails at startup with a message saying so, rather than
  starting with no source configuration.
</Callout>

## Source types [#source-types]

Local backups support eight source types. `script`, `file` and `folder` are local-only,
permanently: they run your commands or read your own disk, neither of which our cloud can or
should do.

### `postgres` [#postgres]

Runs `pg_dump -Fc` (custom format), producing `<database>.dump`.

| Key                    | Required | Notes                                                |
| ---------------------- | -------- | ---------------------------------------------------- |
| `database`             | yes      |                                                      |
| `host`, `port`, `user` | no       | Omitted flags fall back to `pg_dump`'s own defaults  |
| `password`             | no       | Passed as `PGPASSWORD`, never on the command line    |
| `ssl_mode`             | no       | Passed as `PGSSLMODE`, e.g. `require`, `verify-full` |
| `schema`               | no       | Restrict to one schema                               |
| `exclude_tables`       | no       | A list of table names                                |
| `no_owner`             | no       | `true` adds `--no-owner`                             |
| `no_privileges`        | no       | `true` adds `--no-privileges`                        |

```yaml
"<backup-uuid>":
  source_type: postgres
  source:
    database: app
    host: db.internal
    port: 5432
    user: backup
    password: "<password>"
    ssl_mode: require
    exclude_tables: [audit_log, sessions]
    no_owner: true
    no_privileges: true
```

The worker always passes `--no-password`, so `pg_dump` fails fast with a readable error
instead of blocking on an interactive prompt no one will ever answer.

### `mysql` [#mysql]

Runs `mysqldump --single-transaction --quick --routines --events --triggers`, producing
`<database>.sql`.

| Key                    | Required | Notes                                            |
| ---------------------- | -------- | ------------------------------------------------ |
| `database`             | yes      |                                                  |
| `host`, `port`, `user` | no       |                                                  |
| `password`             | no       | Passed as `MYSQL_PWD`, never on the command line |
| `ssl_mode`             | no       | Passed through as `--ssl-mode`                   |
| `tables`               | no       | A list. Restrict to specific tables              |
| `exclude_tables`       | no       | A list. Becomes `--ignore-table=<db>.<table>`    |
| `no_data`              | no       | `true` dumps schema only                         |

Setting &#x2A;*both `tables` and `exclude_tables` is refused at startup.** They express opposite
intentions and the worker will not pick one.

`--single-transaction` takes a consistent snapshot without locking the whole database, which
is the difference between a nightly backup you can run in production and one you cannot.

### `redis` [#redis]

Reads every key with `SCAN` and `DUMP`, producing `<host>.redis.jsonl`. Needs no external tool, and works against managed Redis.

| Key                | Required | Notes                                                                       |
| ------------------ | -------- | --------------------------------------------------------------------------- |
| `host`, `port`     | no       | Also names the artifact; defaults to `redis.redis.jsonl`                    |
| `user`, `password` | no       | Password passed as `REDISCLI_AUTH`                                          |
| `db`               | no       | Database index                                                              |
| `tls`              | no       | `true` adds `--tls`                                                         |
| `insecure`         | no       | `true` skips certificate verification. **Refused unless `tls` is also set** |

A server that refuses the sync exits zero and writes nothing. The worker treats an empty
snapshot as a failure rather than uploading a zero-byte artifact.

### `s3` [#s3]

Copies a bucket or prefix into one tar, producing `<bucket>.tar`. Any S3-compatible store
works: AWS S3, Cloudflare R2, MinIO, Backblaze B2, Wasabi.

| Key                 | Required | Notes                                                          |
| ------------------- | -------- | -------------------------------------------------------------- |
| `bucket`            | yes      |                                                                |
| `access_key_id`     | yes      |                                                                |
| `secret_access_key` | yes      |                                                                |
| `region`            | no       | Defaults to `auto`, which is what R2 and several others want   |
| `endpoint`          | no       | For non-AWS stores. `https://` is added if you omit the scheme |
| `prefix`            | no       | Restrict to one path                                           |
| `use_path_style`    | no       | `true` for MinIO and most self-hosted stores                   |

```yaml
"<backup-uuid>":
  source_type: s3
  source:
    bucket: uploads
    endpoint: minio.internal:9000
    access_key_id: "<key>"
    secret_access_key: "<secret>"
    use_path_style: true
```

An object rewritten while it is being copied fails the run. The tar header is written from
the listing and cannot be corrected afterwards, so the alternative would be a silently
truncated file, which is worse than a failure you can see. An empty bucket or prefix fails
with `EmptySource` rather than producing an empty archive.

### `web` [#web]

Fetches a URL with `curl` and stores the response body.

| Key       | Required | Notes                                                                                     |
| --------- | -------- | ----------------------------------------------------------------------------------------- |
| `url`     | yes      | **Must start with `https://`.** An `http://` URL is refused rather than silently upgraded |
| `method`  | no       | Defaults to curl's own, so `GET`                                                          |
| `headers` | no       | One `Name: value` per line                                                                |

```yaml
"<backup-uuid>":
  source_type: web
  source:
    url: https://api.internal/export
    headers: |
      Authorization: Bearer <token>
      Accept: application/json
```

Redirects are followed, but only to `http` and `https`, so a redirect to `file://` cannot
turn a web backup into a filesystem read. HTTP error statuses fail the run. The artifact is
named from the response's `Content-Disposition`, falling back to the last path segment of the
final URL.

<Callout type="warn">
  Values must not contain a line break. The worker rejects one rather than risk a newline
  becoming a second directive.
</Callout>

### `script` [#script]

Runs a command you own and stores whatever it writes to the path in `$SAVED_OUTPUT`.

| Key           | Required | Notes                                           |
| ------------- | -------- | ----------------------------------------------- |
| `path`        | yes      | Must be executable, unless `interpreter` is set |
| `args`        | no       | A list, one entry per argument                  |
| `interpreter` | no       | Runs `<interpreter> <path> <args...>`           |
| `working_dir` | no       | Directory to run in                             |

```yaml
"<backup-uuid>":
  source_type: script
  source:
    path: /opt/saved/dump-everything.sh
    working_dir: /opt/saved
```

```bash title="/opt/saved/dump-everything.sh"
#!/usr/bin/env bash
set -euo pipefail

pg_dump --format=custom mydb > "$SAVED_OUTPUT"
```

<Callout type="warn">
  **Write to `$SAVED_OUTPUT`, not to stdout.** Anything on stdout or stderr is treated as a
  log and kept for diagnostics. A script that exits zero without writing to that path fails
  with `EmptyOutput`, which is a better outcome than an empty artifact you discover during a
  restore.
</Callout>

### `file` [#file]

Copies one regular file, keeping its name.

| Key               | Required | Notes                                |
| ----------------- | -------- | ------------------------------------ |
| `path`            | yes      |                                      |
| `follow_symlinks` | no       | `true` to back up a symlink's target |

A symlink without `follow_symlinks` is refused rather than quietly backing up the link text.
Directories, sockets, devices and FIFOs are refused with an explanation.

### `folder` [#folder]

Walks a directory tree into a zip, producing `<folder>.zip`.

| Key               | Required | Notes                                                     |
| ----------------- | -------- | --------------------------------------------------------- |
| `path`            | yes      |                                                           |
| `include`         | no       | A list of globs. Empty means everything                   |
| `exclude`         | no       | A list of globs. Applied before `include`                 |
| `follow_symlinks` | no       | `true` resolves links; otherwise they are stored as links |

```yaml
"<backup-uuid>":
  source_type: folder
  source:
    path: /srv/uploads
    exclude: ["*.tmp", "node_modules", ".git"]
```

A pattern matches against the full relative path, the base name, or any single path segment,
so `node_modules` excludes the directory wherever it appears. Excluding a directory skips the
whole subtree. Entries are stored without zip compression, because the encryption step
compresses anyway and doing it twice costs CPU for nothing.

When `follow_symlinks` is set, each resolved target is archived once, so a link cycle cannot
produce an infinite archive.

## Multiple workers on one host [#multiple-workers-on-one-host]

Each worker needs its own working directory, because each needs its own `config.yaml`.

```
/etc/saved/prod/config.yaml     # worker A
/etc/saved/staging/config.yaml  # worker B
```

Run one process per directory. They can share a `local_temp_path`, since staging files are
uniquely named, but giving each its own makes disk pressure easier to attribute.

What you must **not** do is run the same config twice. See
[one process per credential](/docs/workers/lifecycle#one-process-per-credential).

## Next [#next]

<Cards>
  <Card href="/docs/workers/lifecycle" title="Lifecycle" description="What happens after the config is in place." />

  <Card href="/docs/workers/troubleshooting" title="Troubleshooting" description="Every error this file can produce." />
</Cards>
