---
title: "Lifecycle"
description: "How a worker claims work, executes it, and reports back."
url: "https://saved.sh/docs/workers/lifecycle"
---

The worker dials out and polls. Nothing dials in, no port is opened, and no inbound firewall
rule is needed. If the machine can make outbound HTTPS connections, it can run backups.

## Startup [#startup]

Five things happen, in order, and each one fails loudly rather than degrading.

1. **Read the config.** Missing `api_url` or `token` exits immediately.
2. **Prepare the staging directory.** `local_temp_path` is created `0700`. An unwritable path
   exits immediately, before any run can fail halfway through a dump.
3. **Ask who it is.** A single authenticated call to `/v1/auth/worker/whoami` returns the
   worker ID, the workspace namespace, the task queue and the hub address. None of these are
   values you could configure in advance, and none are cached.
4. **Dial the hub** with the same token.
5. **Resolve the dump tools** and log what it found, then start polling.

The whoami response is the only piece of remote configuration in the whole design, and it
carries no secrets and no source data:

| Field                       | Value                                     |
| --------------------------- | ----------------------------------------- |
| `worker_id`                 | This worker's ID                          |
| `name`                      | The name you provisioned it with          |
| `workspace_id`, `namespace` | Your workspace                            |
| `task_queue`                | The queue to poll, which is the worker ID |
| `hub_address`               | The hub to dial                           |

<Callout type="warn">
  Whoami returns no TLS flag, so the worker cannot be told to *stop* using TLS remotely. TLS is
  on by default and only the local config can change that. See
  [connecting to the hub](/docs/workers/configuration#connecting-to-the-hub).
</Callout>

## Routing [#routing]

| Layer      | Value                                                  |
| ---------- | ------------------------------------------------------ |
| Namespace  | Your workspace. A worker can see nothing outside it    |
| Task queue | The worker ID. One queue per worker, not per workspace |

Routing is **per-worker, not per-workspace**. A workspace may run several workers, and each
polls only its own queue, so a run goes to exactly the machine that holds that backup's
credentials. This is why a `local` backup must name a worker, and why changing that name
recreates the schedule.

The worker's connection is authorized to poll its queue, report results, and describe the
namespace. It cannot start work, administer anything, or reach another queue.

## One run, end to end [#one-run-end-to-end]

A schedule fires on our side and puts a task on the worker's queue. The task carries three
identifiers and nothing else: the workspace, the backup, and the run.

**No credentials, no configuration and no data are ever sent to the worker.** It looks the
backup ID up in its own `config.yaml`, because the credentials never left the machine in the
first place, so there is nothing for us to send.

```
Dump<Type>  →  Compress  →  Encrypt  →  Checksum  →  Upload  →  Discard
```

| Step         | Timeout | Attempts | What it does                                          |
| ------------ | ------- | -------- | ----------------------------------------------------- |
| **Dump**     | 6h      | 2        | Runs the source's dump into a scratch file            |
| **Compress** | 2h      | 2        | gzip, if this backup asked for it                     |
| **Encrypt**  | 2h      | 2        | OpenPGP to your public key, into another scratch file |
| **Checksum** | 30m     | 3        | SHA-256 and size **of the uploaded file**             |
| **Upload**   | 1h      | 5        | Presigned PUT, then confirm to the API                |
| **Discard**  | 1m      | 5        | Removes every scratch file the run created            |

**Compress and encrypt are skipped when the backup did not ask for them**, and pass the path
straight through rather than copying it. A backup with neither produces one scratch file, not
three.

Retries back off from 30 seconds for the long steps and 10 seconds for the short ones,
doubling to a ten-minute ceiling.

The dump and encrypt steps get **two attempts, not five**, on purpose: a dump that failed
halfway usually fails the same way again, and each retry re-reads the whole source. Upload
gets five, because a transfer failing is usually the network rather than the data.

Every step emits a **heartbeat every 20 seconds** with its progress, so a six-hour dump is
visibly alive rather than indistinguishable from a hung process. A step that stops
heartbeating for 60 seconds is treated as dead and retried.

### The steps hand each other a path [#the-steps-hand-each-other-a-path]

Each step passes the **location** of a file on your machine to the next one, never the
contents. The bytes never enter our orchestration.

This is also why the file handoff works at all: the worker's task queue is its own ID, so
every step of a run lands on the machine holding the file.

### Discard always runs [#discard-always-runs]

The cleanup step runs in a disconnected context, which means it executes even when the
workflow is failing or has been cancelled. That is exactly when it matters, because a failed
run is the case where a multi-gigabyte scratch file would otherwise be left behind.

If the worker process dies mid-run, the cleanup is not lost. It is queued like any other
step and executes when the worker comes back.

## Upload [#upload]

The worker never holds long-lived storage credentials. It asks the API for a presigned URL
scoped to that one run, uploads directly to object storage, and then confirms.

| Size     | How                                                                                                 |
| -------- | --------------------------------------------------------------------------------------------------- |
| Ordinary | A single presigned `PUT`                                                                            |
| Large    | Multipart. Part URLs are requested 100 at a time, uploaded in order, and completed with their ETags |

A failed multipart upload aborts the upload server-side rather than leaving orphaned parts
behind.

Confirmation is a separate call carrying the checksum, the byte count and the filename. Until
it lands, the run is not finished. Once it does, post-backup takes over on our side: moving
the artifact to permanent storage, indexing it, and applying retention.

<Callout>
  The checksum is taken **after** encryption, so it attests to the bytes we actually store. You
  can verify a downloaded artifact against it without decrypting anything.
</Callout>

## When the worker is offline [#when-the-worker-is-offline]

Runs **queue on the worker's own task queue** and wait. They are not redistributed, because no
other worker has the credentials.

* Bring the worker back within the schedule's window and the queued run executes late.
* Leave it down long enough and runs accumulate, then execute in a burst on return.
* The run is not lost while it waits, and it is not silently dropped.

A worker that is up but never receives work is a different problem, and usually a more
serious one, because nothing about it looks broken. See
[silently idle](/docs/workers/troubleshooting#the-worker-is-up-but-nothing-runs).

## Concurrency [#concurrency]

The worker does not cap how many runs it takes at once. It uses the orchestration SDK's
defaults, which are high enough that in practice the limit is your own schedule: two backups
set to 02:00 will both start at 02:00, on the same machine, competing for the same disk and
the same network.

Two things to plan for:

* **Disk.** Each concurrent run stages a dump and its encrypted copy at the same time. Peak
  usage is roughly twice the sum of the dumps running together, not twice the largest one.
* **Source load.** Two dumps against the same database at the same time is a decision, so make
  it deliberately by staggering the schedules.

If you want genuine isolation between two workloads, run **two workers**, on two hosts, with
two credentials. That is the supported way to divide capacity.

## One process per credential [#one-process-per-credential]

Any number of processes may run the same worker credential, and they will all poll the same
queue. &#x2A;*Do not do this.**

The steps of a run hand each other a path to a local file. With two processes on one queue,
the dump can land on machine A and the upload on machine B, where the path does not exist.
The run fails, and it fails in a way that looks like a storage problem rather than a
configuration one.

<Callout type="warn">
  If the worker list shows more than one instance for a worker you deployed once, find the
  second copy. The usual causes are a container restarted without the old one being stopped, a
  systemd unit racing a manually started process, or the same `config.yaml` copied to a second
  host during a migration.
</Callout>

Scaling out is provisioning a second worker and assigning some backups to it.

## Shutdown and restart [#shutdown-and-restart]

The worker stops on `SIGINT` or `SIGTERM`, finishing what it can and releasing its poll.

A run that was in flight is not lost. Its progress lives in the run record, not in the
process, so when the worker comes back the interrupted step is retried within the attempt
budget above. A dump killed at hour five will start again from the beginning, because there
is no partial dump to resume, which is the strongest argument for scheduling large backups
where a restart is unlikely to land on them.

## Next [#next]

<Cards>
  <Card href="/docs/workers/local-backups" title="What stays local" description="Exactly which bytes leave your network in the steps above." />

  <Card href="/docs/workers/troubleshooting" title="Troubleshooting" description="When one of these steps stops happening." />
</Cards>
