Lifecycle
How a worker claims work, executes it, and reports back.
View as MarkdownThe worker dials out and polls. Nothing dials in, no port is opened, and no inbound firewall rule is needed. If the machine can make outbound HTTPS connections, it can run backups.
Startup
Five things happen, in order, and each one fails loudly rather than degrading.
- Read the config. Missing
api_urlortokenexits immediately. - Prepare the staging directory.
local_temp_pathis created0700. An unwritable path exits immediately, before any run can fail halfway through a dump. - Ask who it is. A single authenticated call to
/v1/auth/worker/whoamireturns the worker ID, the workspace namespace, the task queue and the hub address. None of these are values you could configure in advance, and none are cached. - Dial the hub with the same token.
- Resolve the dump tools and log what it found, then start polling.
The whoami response is the only piece of remote configuration in the whole design, and it carries no secrets and no source data:
| Field | Value |
|---|---|
worker_id | This worker's ID |
name | The name you provisioned it with |
workspace_id, namespace | Your workspace |
task_queue | The queue to poll, which is the worker ID |
hub_address | The hub to dial |
Whoami returns no TLS flag, so the worker cannot be told to stop using TLS remotely. TLS is on by default and only the local config can change that. See connecting to the hub.
Routing
| Layer | Value |
|---|---|
| Namespace | Your workspace. A worker can see nothing outside it |
| Task queue | The worker ID. One queue per worker, not per workspace |
Routing is per-worker, not per-workspace. A workspace may run several workers, and each
polls only its own queue, so a run goes to exactly the machine that holds that backup's
credentials. This is why a local backup must name a worker, and why changing that name
recreates the schedule.
The worker's connection is authorized to poll its queue, report results, and describe the namespace. It cannot start work, administer anything, or reach another queue.
One run, end to end
A schedule fires on our side and puts a task on the worker's queue. The task carries three identifiers and nothing else: the workspace, the backup, and the run.
No credentials, no configuration and no data are ever sent to the worker. It looks the
backup ID up in its own config.yaml, because the credentials never left the machine in the
first place, so there is nothing for us to send.
Dump<Type> → Compress → Encrypt → Checksum → Upload → Discard| Step | Timeout | Attempts | What it does |
|---|---|---|---|
| Dump | 6h | 2 | Runs the source's dump into a scratch file |
| Compress | 2h | 2 | gzip, if this backup asked for it |
| Encrypt | 2h | 2 | OpenPGP to your public key, into another scratch file |
| Checksum | 30m | 3 | SHA-256 and size of the uploaded file |
| Upload | 1h | 5 | Presigned PUT, then confirm to the API |
| Discard | 1m | 5 | Removes every scratch file the run created |
Compress and encrypt are skipped when the backup did not ask for them, and pass the path straight through rather than copying it. A backup with neither produces one scratch file, not three.
Retries back off from 30 seconds for the long steps and 10 seconds for the short ones, doubling to a ten-minute ceiling.
The dump and encrypt steps get two attempts, not five, on purpose: a dump that failed halfway usually fails the same way again, and each retry re-reads the whole source. Upload gets five, because a transfer failing is usually the network rather than the data.
Every step emits a heartbeat every 20 seconds with its progress, so a six-hour dump is visibly alive rather than indistinguishable from a hung process. A step that stops heartbeating for 60 seconds is treated as dead and retried.
The steps hand each other a path
Each step passes the location of a file on your machine to the next one, never the contents. The bytes never enter our orchestration.
This is also why the file handoff works at all: the worker's task queue is its own ID, so every step of a run lands on the machine holding the file.
Discard always runs
The cleanup step runs in a disconnected context, which means it executes even when the workflow is failing or has been cancelled. That is exactly when it matters, because a failed run is the case where a multi-gigabyte scratch file would otherwise be left behind.
If the worker process dies mid-run, the cleanup is not lost. It is queued like any other step and executes when the worker comes back.
Upload
The worker never holds long-lived storage credentials. It asks the API for a presigned URL scoped to that one run, uploads directly to object storage, and then confirms.
| Size | How |
|---|---|
| Ordinary | A single presigned PUT |
| Large | Multipart. Part URLs are requested 100 at a time, uploaded in order, and completed with their ETags |
A failed multipart upload aborts the upload server-side rather than leaving orphaned parts behind.
Confirmation is a separate call carrying the checksum, the byte count and the filename. Until it lands, the run is not finished. Once it does, post-backup takes over on our side: moving the artifact to permanent storage, indexing it, and applying retention.
The checksum is taken after encryption, so it attests to the bytes we actually store. You can verify a downloaded artifact against it without decrypting anything.
When the worker is offline
Runs queue on the worker's own task queue and wait. They are not redistributed, because no other worker has the credentials.
- Bring the worker back within the schedule's window and the queued run executes late.
- Leave it down long enough and runs accumulate, then execute in a burst on return.
- The run is not lost while it waits, and it is not silently dropped.
A worker that is up but never receives work is a different problem, and usually a more serious one, because nothing about it looks broken. See silently idle.
Concurrency
The worker does not cap how many runs it takes at once. It uses the orchestration SDK's defaults, which are high enough that in practice the limit is your own schedule: two backups set to 02:00 will both start at 02:00, on the same machine, competing for the same disk and the same network.
Two things to plan for:
- Disk. Each concurrent run stages a dump and its encrypted copy at the same time. Peak usage is roughly twice the sum of the dumps running together, not twice the largest one.
- Source load. Two dumps against the same database at the same time is a decision, so make it deliberately by staggering the schedules.
If you want genuine isolation between two workloads, run two workers, on two hosts, with two credentials. That is the supported way to divide capacity.
One process per credential
Any number of processes may run the same worker credential, and they will all poll the same queue. Do not do this.
The steps of a run hand each other a path to a local file. With two processes on one queue, the dump can land on machine A and the upload on machine B, where the path does not exist. The run fails, and it fails in a way that looks like a storage problem rather than a configuration one.
If the worker list shows more than one instance for a worker you deployed once, find the
second copy. The usual causes are a container restarted without the old one being stopped, a
systemd unit racing a manually started process, or the same config.yaml copied to a second
host during a migration.
Scaling out is provisioning a second worker and assigning some backups to it.
Shutdown and restart
The worker stops on SIGINT or SIGTERM, finishing what it can and releasing its poll.
A run that was in flight is not lost. Its progress lives in the run record, not in the process, so when the worker comes back the interrupted step is retried within the attempt budget above. A dump killed at hour five will start again from the beginning, because there is no partial dump to resume, which is the strongest argument for scheduling large backups where a restart is unlikely to land on them.