Upgrades
Keeping workers current without breaking running backups.
View as MarkdownUpgrading a worker is replacing a binary and restarting a process. There is no migration, no re-registration, and no state on the worker to carry forward: the credential, the worker ID and the backups assigned to it all live outside the process.
Releases
| Channel | Where | What |
|---|---|---|
| Binaries | GitHub releases | Six targets plus SHA256SUMS |
| Container | ghcr.io/savedhq/local-worker | vX.Y.Z, latest, sha-<short> |
Binaries are published for Linux, macOS and Windows on amd64 and arm64. The container
image is currently built for linux/arm64; on an amd64 host, use the released binary or
build the image yourself.
Pin a version. latest follows the tip of the default branch, which is the right choice
for a machine you are actively testing against and the wrong one for the host that holds your
production backups.
What has to stay compatible
A worker talks to us over three contracts, and only the first one can break quietly.
| Contract | What breaks | How it shows up |
|---|---|---|
| Workflow and activity names | A worker registering names we do not schedule, or missing one we do | Runs never start. The worker looks healthy and idle |
| The whoami response | A worker that cannot parse it | Fails at startup, loudly |
| The run API | Upload or confirm calls rejected | Runs fail at the upload step, loudly |
The first row is the reason to keep workers reasonably current: a worker that does not register a source type's workflow does not fail, it simply never receives that work. A schedule fires into nothing, and nothing reports an error. When you add a backup of a source type you have not used before, check that the run actually starts rather than assuming silence means success.
Two of these have bitten us before, in both directions: a name registered that nothing scheduled, and a name scheduled that nothing registered. Neither produced an alert, which is why the names are treated as a wire contract rather than an implementation detail.
Upgrading in place
sudo systemctl stop saved-worker
VERSION=v1.1.0
curl -fsSLO "https://github.com/savedhq/local-worker/releases/download/${VERSION}/local-worker-linux-amd64"
curl -fsSLO "https://github.com/savedhq/local-worker/releases/download/${VERSION}/SHA256SUMS"
grep 'local-worker-linux-amd64$' SHA256SUMS | sha256sum -c -
sudo install -m 0755 local-worker-linux-amd64 /usr/local/bin/local-worker
sudo systemctl start saved-workerFor the container, change the tag and recreate:
docker rm -f saved-worker
docker run -d --name saved-worker \
--restart unless-stopped \
-v /etc/saved/config.yaml:/config.yaml:ro \
-v saved-scratch:/tmp/saved \
ghcr.io/savedhq/local-worker:v1.1.0Check the startup log before walking away. A worker that failed to parse a changed config exits rather than running degraded, so a silent process list is not evidence of success.
What happens to a run in flight
Stopping the worker stops it mid-step. Nothing is lost, because the run's progress lives in the run record rather than in the process.
| Step interrupted | On restart |
|---|---|
| Dump | Retried from the beginning. There is no partial dump to resume |
| Encrypt | Retried from the beginning |
| Checksum | Retried |
| Upload | Retried, including a multipart upload, which restarts cleanly |
| Any | Scratch files are cleaned up by the queued cleanup step |
Each step has a limited attempt budget, so an upgrade that lands in the middle of the second attempt of a six-hour dump can exhaust it and fail the run. Two practical consequences:
- Upgrade outside the backup window. Not because the design is fragile, but because re-running a large dump costs real time on the source.
- A failed run is not a lost backup. The next scheduled run produces a fresh artifact, and the previous artifact is untouched.
Scratch cleanup survives the restart. It is queued like any other step and runs when the
worker returns, so an interrupted upgrade does not leave multi-gigabyte files behind
permanently. If the worker is never coming back, remove local_temp_path by hand.
Knowing what each worker runs
We do not collect the worker's version. The API knows a worker's name, ID, how many instances are polling and when it was last seen. It does not know which binary those instances are.
sctl worker listTo find the version, read it from the host:
# systemd
journalctl -u saved-worker | grep 'local-worker starting'
# container
docker inspect --format '{{.Config.Image}}' saved-workerIf you run more than a couple of workers, record the version in your own configuration management rather than discovering it during an incident. Worker versions are the one part of this system we cannot tell you about.
Rolling back
Roll back the same way you rolled forward: stop, install the previous binary or tag, start. Nothing in a run's history is version-specific, and artifacts produced by any version are readable by any other, because the format is OpenPGP and a checksum rather than anything of ours. See Artifact format.