Here is a pg_dump to S3 script that works. Paste it, schedule it, and you have
backups tonight.
#!/usr/bin/env bash
set -euo pipefail
STAMP="$(date -u +%Y-%m-%dT%H-%M-%SZ)"
KEY="postgres/${PGDATABASE}/${STAMP}.dump.gz"
pg_dump --format=custom --no-owner --no-privileges \
| gzip -9 \
| aws s3 cp - "s3://${BUCKET}/${KEY}" \
--expected-size 5000000000
echo "wrote s3://${BUCKET}/${KEY}"Three details in there are worth keeping whatever you do next. set -euo pipefail means a failure in the middle of the pipe actually fails the script,
which the default shell behaviour does not. --format=custom gives you
selective restore with pg_restore later. --expected-size lets the AWS CLI
choose a part size big enough to stream past the 10,000 part limit, which is the
thing that silently breaks the first time a database crosses about 80 GB.
That script is not wrong. It is just much smaller than the problem.
The six things it cannot see
1. A dump that succeeded and is useless. pg_dump exits zero having produced
a valid dump of a replica that stopped replicating six weeks ago. Nothing in the
pipeline compares last night's size to the night before, which is the cheapest
possible check and catches an enormous share of real incidents.
2. Silence as a success signal. Cron mails output to a local mailbox nobody reads. If the box is down, the job does not fail. It simply does not run, and absence of failure looks identical to success. You find out during a restore.
3. The run that outlasts its window. A nightly job that starts taking 25 hours now has two copies of itself running, competing for the same source and writing to keys that differ by a timestamp. Neither is complete.
4. Retention that is either absent or a footgun. A bucket lifecycle rule deletes by age with no idea whether a newer backup ever succeeded. If dumps have been failing for 40 days and your rule expires at 30, the rule deletes your last good copy on schedule.
5. Credentials at rest on the box. The script needs database credentials and write credentials to the bucket, sitting in an environment file on a host whose access list has drifted since it was provisioned.
6. The blast radius. The bucket is usually in the same cloud account as the database. One compromised set of credentials, or one enthusiastic Terraform destroy, takes both.
6
Failure modes above the script cannot detect
0
Of them produce a non-zero exit code
1
Test that finds most of them
Point it at a replica, and then watch the replica
One change is worth making before any of the above, and it is free if you already have a read replica: dump the replica, not the primary.
A dump is a long read holding a snapshot open for its whole duration. On the
primary that means competing with production for I/O, blocking exclusive DDL, and
holding back VACUUM so dead tuples pile up for however many hours it runs. On a
small database nobody notices. On a large one it is the reason the backup window
is a thing people schedule around.
Then it introduces a failure mode worse than the six above, and almost nobody instruments it.
A stalled replica backs up stale data, silently. Replication stops. Nothing alerts, because nothing is watching. Every nightly dump afterwards is a faithful, correctly-sized, perfectly restorable copy of a database frozen at the moment replication broke. You find out when you restore and the data is six weeks old.
One more thing specific to Postgres. A long pg_dump against a hot standby can be
killed mid-run by WAL replay with canceling statement due to conflict with recovery. hot_standby_feedback = on fixes it at the cost of bloat back on the
primary you were trying to protect. max_standby_streaming_delay = -1 fixes it by
letting the standby fall behind while the dump runs, which is usually the right
trade for a node whose job is backups.
What fixing each one actually requires
The uncomfortable part is that every fix is small, and there are a lot of them.
| Failure | The fix |
|---|---|
| Useless dump | Record size and duration per run, alert on deviation |
| Stale replica | Record replay lag per run, alert when the size stops moving |
| Silence | Push a heartbeat outward, alert on the absence of a run |
| Overlap | A scheduler that knows whether the previous run finished |
| Retention | Expiry that is aware of whether newer copies exist |
| Credentials | A secret store and short-lived credentials, not an env file |
| Blast radius | A destination in a different account with different credentials |
None of these is hard. Together they are a system that somebody on your team now owns, and it is nobody's favourite work.
What we do instead
We run the same dump. The difference is what surrounds it.
Every run produces a record: what ran, how long it took, how many bytes it produced, where each copy landed, and what failed if anything did. Runs are durable, so a worker dying halfway through does not quietly skip a night. Retention is expressed as a policy rather than a lifecycle rule, so expiry knows about your other copies. Credentials live in a secret store and are read at run time by the worker that needs them.
The bytes go to a bucket you own, compressed and encrypted before they leave the machine that produced them.
The honest summary
If you have one database and one engineer who remembers to check, the script above is genuinely fine. The moment you have five sources, two clouds, or an auditor asking when the last successful restore was, you are building the second column of that table whether you meant to or not.