---
title: "Troubleshooting"
description: "What to look at when runs stop happening."
url: "https://saved.sh/docs/workers/troubleshooting"
---

Almost every worker problem is one of four things: the config is wrong, the key is no longer
valid, the host cannot reach the hub, or the source itself is refusing. The startup log tells
you which within the first two seconds.

## What healthy looks like [#what-healthy-looks-like]

The worker logs to **stderr** in `key=value` form. A good start looks like this:

```
level=INFO msg="external tool resolved" tool=curl path=/usr/bin/curl
level=INFO msg="external tool resolved" tool=pg_dump path=/usr/bin/pg_dump
level=INFO msg="hub dial resolved" address=hub.saved.sh:443 namespace=<workspace-id> tls_enabled=true task_queue=<worker-id>
level=INFO msg="local-worker starting" api_url=https://api.saved.sh backups=3 temp_path=/tmp/saved
```

Four things to check in that output, in order:

| Line                     | What to confirm                                                          |
| ------------------------ | ------------------------------------------------------------------------ |
| `external tool resolved` | Every tool your source types need is listed                              |
| `hub dial resolved`      | `tls_enabled=true` against `hub.saved.sh`, and a `task_queue` is present |
| `local-worker starting`  | `backups=` matches the number of backups assigned to this worker         |
| Nothing after it         | Silence is correct. A polling worker logs nothing until work arrives     |

`backups=0` is the single most useful early warning: the worker is connected and will fail
every run it receives.

Set `debug: true` for per-poll detail. It is verbose enough that you will not want it left on.

## Startup failures [#startup-failures]

The worker exits rather than running degraded. Each of these is the last line before it does.

| Message                                                        | Cause                                             | Fix                                                                                                                               |
| -------------------------------------------------------------- | ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `missing required config: api_url, token`                      | Empty or missing keys                             | Check the file is where the worker looks. A missing file is tolerated silently, and then this is what you get                     |
| `read config "config.yaml": ...`                               | The file exists but is not valid YAML             | Check the indentation. A tab where YAML wants spaces is the usual one                                                             |
| `parse backups in config.yaml: ...`                            | A `backups:` entry has the wrong shape            | Each entry needs `source_type` and `source`                                                                                       |
| `backup <id> has no source_type`                               | An entry is missing its type                      | Add it                                                                                                                            |
| `backup <id> still uses \`credentials:\`\`                     | The key was renamed to `source:`                  | Rename it. The fields inside are unchanged                                                                                        |
| `backup <id>: <type> source: ... has invalid keys: <key>`      | A field this source type does not have            | Check the spelling against [Source types](/docs/workers/configuration#source-types). Unknown keys are refused rather than ignored |
| `backup <id>: <type> source: <field> is required`              | A required field is missing                       | [Source types](/docs/workers/configuration#source-types) lists which are required                                                 |
| `backup <id>: encryption.public_key: ...`                      | The key file is missing, or holds no armoured key | A relative path resolves against the config file's directory                                                                      |
| `temp path unusable`                                           | `local_temp_path` cannot be created or written    | Check ownership. In a container, the worker runs as uid `65532`                                                                   |
| `fetch worker.whoami: worker.whoami returned 401 Unauthorized` | The key is wrong or has been revoked              | The key was rotated or the worker deleted. Paste the current key                                                                  |
| `fetch worker.whoami: worker.whoami returned 404`              | The key is valid but is not a worker key          | You have pasted an API key. Provision a worker and use its key                                                                    |
| `no task queue: worker.whoami did not return a worker id`      | The credential is not bound to a worker           | As above                                                                                                                          |
| `dial hub hub.saved.sh:443: ...`                               | Cannot reach the hub                              | See below                                                                                                                         |

### The dial hangs or times out [#the-dial-hangs-or-times-out]

The address, namespace and task queue all come from the API at startup, so a hang is almost
never the endpoint being wrong. Confirm what was actually resolved with the
`hub dial resolved` log line, then work through these:

* **Egress filtering.** The worker needs outbound access to `hub.saved.sh` and to your
  `api_url`. Nothing needs to reach the worker.
* **An inspecting proxy.** If TLS is terminated and re-signed on the way out, set
  `hub_tls_ca_path` to the proxy's CA. Leave it empty otherwise; empty means the system trust
  store, which is what a public certificate needs.
* **Clock skew.** A host more than a few minutes off will fail certificate validation. Check
  `timedatectl` before assuming anything more exotic.

## Run failures [#run-failures]

These come back on the run itself, in the dashboard or `sctl run list`. The ones marked
non-retryable fail immediately and will keep failing until you change something, which is the
point: retrying a misconfiguration wastes hours.

| Error                | Meaning                                                                                                                                                              |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `ConfigDrift`        | A run fired for a backup with no entry in this worker's `backups:`. Add the entry, or remove the schedule                                                            |
| `SourceTypeMismatch` | The entry's `source_type` disagrees with the run that fired. The backup was changed on one side only                                                                 |
| `MissingTool`        | The source type needs `pg_dump`, `mysqldump` or `curl` and it is not on `PATH`. Install it or set `tools.<name>`                                                     |
| `InvalidSource`      | A required credential is missing, or a path is not what it claims: a symlink without `follow_symlinks`, a directory given to `file`, a script that is not executable |
| `EmptyOutput`        | A `script` source exited zero without writing to `$SAVED_OUTPUT`. Anything on stdout is treated as a log, not as data                                                |
| `EmptySource`        | An `s3` bucket or prefix has no objects                                                                                                                              |

Failures that are worth retrying, and are retried automatically:

| Symptom                                                    | What it usually is                                                                                           |
| ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| `pg_dump app: FATAL: password authentication failed`       | The credential in `config.yaml`. The message is the tool's own                                               |
| `snapshot redis: connect to redis at ...`                  | The host, port or password is wrong, or the server is unreachable                                            |
| `<key> changed while being copied: listed N bytes, read M` | An S3 object was rewritten mid-run. Rerun, or back up a prefix that is not being actively written            |
| `storage PUT returned 403`                                 | The presigned URL expired, usually after the upload took longer than the window. Retries request a fresh URL |
| Disk errors during dump or encrypt                         | `local_temp_path` filled up. Budget roughly twice the dump size, per concurrent run                          |

<Callout>
  Error text is the failing tool's own stderr, truncated to 2000 characters, because "access
  denied for user [backup@db.internal](mailto:backup@db.internal)" is worth more than "exit status 2". If a message looks
  cut off, that is why. Reproduce the command on the host for the full output.
</Callout>

## The worker is up but nothing runs [#the-worker-is-up-but-nothing-runs]

This is the worst failure shape available, because everything looks healthy. Work through it
in this order:

1. **Is the backup assigned to this worker?** A `local` backup names exactly one worker, and
   its runs go only to that worker's queue. A backup pointing at a worker you decommissioned
   queues forever.
2. **Does `backups=` in the startup log match?** If the worker knows about zero backups, the
   config's keys are wrong. They are **backup IDs**, not names.
3. **Is the schedule actually set?** Check `sctl backup list` for the backup's schedule, and
   `sctl run list --backup <name>` for whether runs are being created at all.
4. **Is more than one instance polling?** Check the worker list. Two processes on one
   credential produce runs that fail at a step that cannot find its file.
5. **Is the source type one this worker's version supports?** A worker that does not register
   a source type's workflow receives nothing for it and reports no error. Upgrade and retry.

Trigger a run by hand rather than waiting for the schedule, which separates "the worker is
broken" from "the schedule never fired":

```bash
sctl backup trigger $BACKUP_ID
sctl run list --backup $BACKUP_ID
```

## Failed runs versus missing runs [#failed-runs-versus-missing-runs]

They are different problems with different evidence.

|                     | Failed                                                | Missing                                          |
| ------------------- | ----------------------------------------------------- | ------------------------------------------------ |
| A run record exists | Yes, with an error                                    | No                                               |
| Where to look       | The run's error text, and the worker log at that time | The schedule, and the backup's worker assignment |
| Usual cause         | The source, the config, or the host                   | Routing: no worker, wrong worker, or no schedule |

A worker that was **offline** at schedule time produces neither. Its runs queue on its own
task queue and execute when it returns, late rather than lost. If you brought a worker back
and a burst of runs started, that is the expected behaviour and not a fault.

## Checking the pieces individually [#checking-the-pieces-individually]

Confirm the credential is live, without involving the worker at all:

```bash
curl -fsS https://api.saved.sh/v1/auth/worker/whoami \
  -H "Authorization: Bearer $TOKEN"
```

A `200` with a `task_queue` means the key is valid and bound to a worker. A `401` means it is
revoked or wrong. A `403` means it is not a worker key.

Confirm the dump works, as the user the worker runs as:

```bash
sudo -u saved PGPASSWORD=... pg_dump -Fc -f /tmp/test.dump -h db.internal -U backup app
```

Most "the worker is broken" reports are a source credential that works for you interactively
and not for the service user.

## Writing in [#writing-in]

Support can act immediately with these, and not much without them:

* The **worker ID** and the **run ID** of a failing run.
* The worker's startup log: the four lines above are enough, and they contain no secrets.
* The error text from the run, in full.
* The worker version, which we cannot see. See
  [knowing what each worker runs](/docs/workers/upgrades#knowing-what-each-worker-runs).
* Whether it ever worked, and what changed if it did.

<Callout type="warn">
  Do not send us your `config.yaml`. It contains your source credentials, and we have no use
  for them. If a config question is genuinely the issue, redact every value and send the shape.
</Callout>

See [Support](/docs/reference/support).
