Troubleshooting
A pod that will not start, or starts and does nothing.
View as MarkdownAlmost every failure here is one of four things: the pod cannot read its config, cannot write its scratch volume, cannot run the image, or is running perfectly and has nothing to do. The first log lines say which.
kubectl -n saved logs deploy/saved-worker
kubectl -n saved describe pod -l app.kubernetes.io/name=saved-workerThe general worker troubleshooting page covers everything that is not cluster-specific, including what a healthy start looks like and what each run error means. This page is the parts that are.
The pod will not start
| Symptom | Cause | Fix |
|---|---|---|
no matching manifest for linux/amd64 | The tag has no image for the node's architecture | docker buildx imagetools inspect the tag, then pin nodeSelector or build your own |
CrashLoopBackOff, log says missing required config: api_url, token | The worker looked in the wrong directory, or the Secret key is not config.yaml | workingDir must match the volume's mountPath, and the Secret key must be exactly config.yaml |
CrashLoopBackOff, log says temp path unusable | The scratch volume is not writable by uid 65532 | Set fsGroup: 65532 on the pod, and local_temp_path to the mount path |
CreateContainerConfigError | The Secret does not exist in that namespace | kubectl -n saved get secret saved-worker-config |
| Restarts every 30 seconds, no error in the log | A readiness or liveness probe | Remove it. The worker serves no HTTP and opens no port |
missing required config usually means the file is not there at all. A missing
config.yaml is tolerated silently, so the worker gets as far as validating an empty config
before it exits. kubectl -n saved exec deploy/saved-worker -- ls -l /etc/saved settles it,
though a crash-looping pod will not let you.
The pod is running and nothing happens
Silence is correct. A worker that has connected logs nothing until work arrives. Check the last startup lines rather than waiting for output.
| What to check | How |
|---|---|
| It declares the backups you expect | backups= in the local-worker starting line. backups=0 means it will fail every run it receives |
| It is polling | sctl worker list, which reads presence from the live connection |
| It resolved the tools | The external tool resolved lines, one per tool the image ships |
If backups=0, the Secret holds a config with no backups: map. That is the most common
cluster mistake, because a config assembled from environment variables cannot carry one.
Runs fail after the pod is healthy
| Symptom | Cause |
|---|---|
ConfigDrift on every run | The backup ID in the dashboard is not in this worker's backups: |
| Disk errors partway through a dump | The emptyDir filled. Budget roughly twice the dump size, and set a sizeLimit you can afford |
| The dump cannot reach the database | Service DNS, or a NetworkPolicy with no egress rule. Test with kubectl -n saved exec deploy/saved-worker -- curl -sv telnet://postgres.default.svc.cluster.local:5432 |
A folder source sees nothing | The claim is not mounted into the worker pod, or is mounted at a different path than source.path |
Runs break in ways that make no sense
Check the replica count first.
kubectl -n saved get deploy saved-worker -o jsonpath='{.spec.replicas} {.spec.strategy.type}'The answer must be 1 Recreate. Two processes sharing one worker credential both poll the
same queue, and a run's steps hand each other a path to a file on local disk, so the second
pod looks for a file that is not there. A rolling update does this briefly on every deploy,
which makes it look intermittent rather than structural.
Why.
A rotated key did not take effect
Rotation revokes the old key immediately, so the pod must restart.
kubectl -n saved rollout restart deploy/saved-workerIf it still uses the old key, the Secret is mounted with subPath. A subPath mount never
receives updates. Mount the Secret as a directory and set workingDir to it instead. See
Install.
The operator
Not yet available
There is nothing to troubleshoot yet. A LocalWorker resource is accepted by the API server
and then nothing acts on it, so an empty status is the expected result rather than a
symptom.