Operations
Upgrading, watching, and what breaks when something is down.
View as MarkdownRunning the worker in a cluster is running one stateless pod that dials out. There is nothing to fail over, no leader to elect and no data on the pod worth preserving.
Upgrading the worker
Change the tag and let the Deployment recreate the pod.
kubectl -n saved set image deploy/saved-worker worker=ghcr.io/savedhq/local-worker:v0.2.0
kubectl -n saved rollout status deploy/saved-workerstrategy: Recreate, not RollingUpdate. A rolling update briefly runs two pods on one
worker credential, and a run in flight during that window can have its steps split across
them. The old pod stopping before the new one starts is the correct behaviour here, not a
compromise. Why.
A run in flight when the pod stops is not lost. It is queued on the worker's own task queue and resumes when the new pod connects. See what happens to a run in flight.
Keeping workers current
The reason to stay reasonably current is quiet: a worker that does not register a source type's workflow does not fail, it never receives that work. A schedule fires into nothing and nothing reports an error. When you add a backup of a source type this cluster has not run before, check that the first run actually starts.
Pin a tag. latest follows the tip of the default branch, which is the right choice for a
cluster you are testing against and the wrong one for the one holding your production
backups.
Watching it
There are no metrics and no /healthz. The worker opens no port at all, which is why the
Deployment has no probes and no Service.
| Question | Where the answer is |
|---|---|
| Is the pod up | kubectl -n saved get pods |
| Is the worker polling | sctl worker list, read from the live connection rather than a stored heartbeat |
| How many backups does it serve | The local-worker starting line in the pod log |
| Did a run succeed | sctl run list, or the dashboard. Runs are recorded on our side |
Alert on runs, not on the pod. A pod can be up, connected, and serving zero backups; a run that did not happen is the thing that costs you. See Notifications.
The version a pod is running is not something we collect. Read it from the image tag:
kubectl -n saved get deploy saved-worker -o jsonpath='{.spec.template.spec.containers[0].image}'Capacity
One pod runs one worker. Two workers means two Deployments, two Secrets and two provisioned credentials, never two replicas.
Size the scratch volume for the largest dump plus its encrypted copy, and multiply by the number of runs that can overlap. Compression happens before encryption, so a compressed backup needs less at the upload end and the same at the dump end.
What breaks when something is down
| Down | Effect |
|---|---|
| The worker pod | Runs queue on its task queue and execute when it comes back. Nothing is redistributed, because no other worker has the credentials |
| The cluster's egress | Same as the pod being down. The worker retries |
| Our API | Scheduled runs do not fire. Nothing is lost, and nothing on your side needs doing |
The operator
Not yet available
This section describes the intended design.
The operator will own the Deployment, so upgrading a worker becomes editing
spec.image on a LocalWorker, and the Recreate rule stops
being something you have to remember.
Upgrading the operator will be independent of the workers it manages, and a worker keeps running while the controller is down: the controller reconciles desired state, it is not in the path of a backup. Nothing about a run passes through it.