Troubleshooting
What to look at when runs stop happening.
View as MarkdownAlmost every worker problem is one of four things: the config is wrong, the key is no longer valid, the host cannot reach the hub, or the source itself is refusing. The startup log tells you which within the first two seconds.
What healthy looks like
The worker logs to stderr in key=value form. A good start looks like this:
level=INFO msg="external tool resolved" tool=curl path=/usr/bin/curl
level=INFO msg="external tool resolved" tool=pg_dump path=/usr/bin/pg_dump
level=INFO msg="hub dial resolved" address=hub.saved.sh:443 namespace=<workspace-id> tls_enabled=true task_queue=<worker-id>
level=INFO msg="local-worker starting" api_url=https://api.saved.sh backups=3 temp_path=/tmp/savedFour things to check in that output, in order:
| Line | What to confirm |
|---|---|
external tool resolved | Every tool your source types need is listed |
hub dial resolved | tls_enabled=true against hub.saved.sh, and a task_queue is present |
local-worker starting | backups= matches the number of backups assigned to this worker |
| Nothing after it | Silence is correct. A polling worker logs nothing until work arrives |
backups=0 is the single most useful early warning: the worker is connected and will fail
every run it receives.
Set debug: true for per-poll detail. It is verbose enough that you will not want it left on.
Startup failures
The worker exits rather than running degraded. Each of these is the last line before it does.
| Message | Cause | Fix |
|---|---|---|
missing required config: api_url, token | Empty or missing keys | Check the file is where the worker looks. A missing file is tolerated silently, and then this is what you get |
read config "config.yaml": ... | The file exists but is not valid YAML | Check the indentation. A tab where YAML wants spaces is the usual one |
parse backups in config.yaml: ... | A backups: entry has the wrong shape | Each entry needs source_type and source |
backup <id> has no source_type | An entry is missing its type | Add it |
backup <id> still uses \credentials:`` | The key was renamed to source: | Rename it. The fields inside are unchanged |
backup <id>: <type> source: ... has invalid keys: <key> | A field this source type does not have | Check the spelling against Source types. Unknown keys are refused rather than ignored |
backup <id>: <type> source: <field> is required | A required field is missing | Source types lists which are required |
backup <id>: encryption.public_key: ... | The key file is missing, or holds no armoured key | A relative path resolves against the config file's directory |
temp path unusable | local_temp_path cannot be created or written | Check ownership. In a container, the worker runs as uid 65532 |
fetch worker.whoami: worker.whoami returned 401 Unauthorized | The key is wrong or has been revoked | The key was rotated or the worker deleted. Paste the current key |
fetch worker.whoami: worker.whoami returned 404 | The key is valid but is not a worker key | You have pasted an API key. Provision a worker and use its key |
no task queue: worker.whoami did not return a worker id | The credential is not bound to a worker | As above |
dial hub hub.saved.sh:443: ... | Cannot reach the hub | See below |
The dial hangs or times out
The address, namespace and task queue all come from the API at startup, so a hang is almost
never the endpoint being wrong. Confirm what was actually resolved with the
hub dial resolved log line, then work through these:
- Egress filtering. The worker needs outbound access to
hub.saved.shand to yourapi_url. Nothing needs to reach the worker. - An inspecting proxy. If TLS is terminated and re-signed on the way out, set
hub_tls_ca_pathto the proxy's CA. Leave it empty otherwise; empty means the system trust store, which is what a public certificate needs. - Clock skew. A host more than a few minutes off will fail certificate validation. Check
timedatectlbefore assuming anything more exotic.
Run failures
These come back on the run itself, in the dashboard or sctl run list. The ones marked
non-retryable fail immediately and will keep failing until you change something, which is the
point: retrying a misconfiguration wastes hours.
| Error | Meaning |
|---|---|
ConfigDrift | A run fired for a backup with no entry in this worker's backups:. Add the entry, or remove the schedule |
SourceTypeMismatch | The entry's source_type disagrees with the run that fired. The backup was changed on one side only |
MissingTool | The source type needs pg_dump, mysqldump or curl and it is not on PATH. Install it or set tools.<name> |
InvalidSource | A required credential is missing, or a path is not what it claims: a symlink without follow_symlinks, a directory given to file, a script that is not executable |
EmptyOutput | A script source exited zero without writing to $SAVED_OUTPUT. Anything on stdout is treated as a log, not as data |
EmptySource | An s3 bucket or prefix has no objects |
Failures that are worth retrying, and are retried automatically:
| Symptom | What it usually is |
|---|---|
pg_dump app: FATAL: password authentication failed | The credential in config.yaml. The message is the tool's own |
snapshot redis: connect to redis at ... | The host, port or password is wrong, or the server is unreachable |
<key> changed while being copied: listed N bytes, read M | An S3 object was rewritten mid-run. Rerun, or back up a prefix that is not being actively written |
storage PUT returned 403 | The presigned URL expired, usually after the upload took longer than the window. Retries request a fresh URL |
| Disk errors during dump or encrypt | local_temp_path filled up. Budget roughly twice the dump size, per concurrent run |
Error text is the failing tool's own stderr, truncated to 2000 characters, because "access denied for user backup@db.internal" is worth more than "exit status 2". If a message looks cut off, that is why. Reproduce the command on the host for the full output.
The worker is up but nothing runs
This is the worst failure shape available, because everything looks healthy. Work through it in this order:
- Is the backup assigned to this worker? A
localbackup names exactly one worker, and its runs go only to that worker's queue. A backup pointing at a worker you decommissioned queues forever. - Does
backups=in the startup log match? If the worker knows about zero backups, the config's keys are wrong. They are backup IDs, not names. - Is the schedule actually set? Check
sctl backup listfor the backup's schedule, andsctl run list --backup <name>for whether runs are being created at all. - Is more than one instance polling? Check the worker list. Two processes on one credential produce runs that fail at a step that cannot find its file.
- Is the source type one this worker's version supports? A worker that does not register a source type's workflow receives nothing for it and reports no error. Upgrade and retry.
Trigger a run by hand rather than waiting for the schedule, which separates "the worker is broken" from "the schedule never fired":
sctl backup trigger $BACKUP_ID
sctl run list --backup $BACKUP_IDFailed runs versus missing runs
They are different problems with different evidence.
| Failed | Missing | |
|---|---|---|
| A run record exists | Yes, with an error | No |
| Where to look | The run's error text, and the worker log at that time | The schedule, and the backup's worker assignment |
| Usual cause | The source, the config, or the host | Routing: no worker, wrong worker, or no schedule |
A worker that was offline at schedule time produces neither. Its runs queue on its own task queue and execute when it returns, late rather than lost. If you brought a worker back and a burst of runs started, that is the expected behaviour and not a fault.
Checking the pieces individually
Confirm the credential is live, without involving the worker at all:
curl -fsS https://api.saved.sh/v1/auth/worker/whoami \
-H "Authorization: Bearer $TOKEN"A 200 with a task_queue means the key is valid and bound to a worker. A 401 means it is
revoked or wrong. A 403 means it is not a worker key.
Confirm the dump works, as the user the worker runs as:
sudo -u saved PGPASSWORD=... pg_dump -Fc -f /tmp/test.dump -h db.internal -U backup appMost "the worker is broken" reports are a source credential that works for you interactively and not for the service user.
Writing in
Support can act immediately with these, and not much without them:
- The worker ID and the run ID of a failing run.
- The worker's startup log: the four lines above are enough, and they contain no secrets.
- The error text from the run, in full.
- The worker version, which we cannot see. See knowing what each worker runs.
- Whether it ever worked, and what changed if it did.
Do not send us your config.yaml. It contains your source credentials, and we have no use
for them. If a config question is genuinely the issue, redact every value and send the shape.
See Support.