saved.sh
DocsPricingDownloadBlog

By source

  • Database backupsPostgres today, on any of the three paths.
  • Files and foldersDirectories on your own hardware. Local path only.
  • Anything you can scriptThe escape hatch, with no reduced guarantees.

By how it runs

  • Backups you driveCLI or API, driven by whatever you already run.
  • Backups on your infrastructureOur worker, your hardware. We never hold the credential.
  • Fully managed backupsWe execute and orchestrate. Nothing to host.
  • Compliance and custodyYour bucket, your account, our orchestration.

Understand it

  • How it worksThree paths, one artifact lifecycle
  • SecurityWhat we can and cannot see
  • CompareAgainst snapshots, cloud-native and scripts
Get started
saved.sh

External backups for the systems a business actually runs on.

A product of reops.

Product

  • Documentation
  • Solutions
  • Compare
  • How it works
  • Security
  • Pricing
  • Download

Developers

  • CLI
  • REST API
  • Recover

Company

  • Blog
  • Privacy
  • Terms

© 2026 saved.sh

Your data survives what holds it.

All postsEngineering

How to back up PostgreSQL to S3, and what the cron job leaves out

A working pg_dump to S3 script you can paste today, followed by the six failure modes it cannot see. The script is not wrong. It is just much smaller than the problem.

8 August 2026·saved.shView as Markdown

Here is a pg_dump to S3 script that works. Paste it, schedule it, and you have backups tonight.

#!/usr/bin/env bash
set -euo pipefail

STAMP="$(date -u +%Y-%m-%dT%H-%M-%SZ)"
KEY="postgres/${PGDATABASE}/${STAMP}.dump.gz"

pg_dump --format=custom --no-owner --no-privileges \
  | gzip -9 \
  | aws s3 cp - "s3://${BUCKET}/${KEY}" \
      --expected-size 5000000000

echo "wrote s3://${BUCKET}/${KEY}"

Three details in there are worth keeping whatever you do next. set -euo pipefail means a failure in the middle of the pipe actually fails the script, which the default shell behaviour does not. --format=custom gives you selective restore with pg_restore later. --expected-size lets the AWS CLI choose a part size big enough to stream past the 10,000 part limit, which is the thing that silently breaks the first time a database crosses about 80 GB.

That script is not wrong. It is just much smaller than the problem.

The six things it cannot see

A CRON JOBworker restartsthe run is gone, and nothing says soA DURABLE RUNresumes from the step it reachedcompletes
Each of these is invisible from inside the script. The script exits zero and the bucket has an object in it.

1. A dump that succeeded and is useless. pg_dump exits zero having produced a valid dump of a replica that stopped replicating six weeks ago. Nothing in the pipeline compares last night's size to the night before, which is the cheapest possible check and catches an enormous share of real incidents.

2. Silence as a success signal. Cron mails output to a local mailbox nobody reads. If the box is down, the job does not fail. It simply does not run, and absence of failure looks identical to success. You find out during a restore.

3. The run that outlasts its window. A nightly job that starts taking 25 hours now has two copies of itself running, competing for the same source and writing to keys that differ by a timestamp. Neither is complete.

4. Retention that is either absent or a footgun. A bucket lifecycle rule deletes by age with no idea whether a newer backup ever succeeded. If dumps have been failing for 40 days and your rule expires at 30, the rule deletes your last good copy on schedule.

5. Credentials at rest on the box. The script needs database credentials and write credentials to the bucket, sitting in an environment file on a host whose access list has drifted since it was provisioned.

6. The blast radius. The bucket is usually in the same cloud account as the database. One compromised set of credentials, or one enthusiastic Terraform destroy, takes both.

6

Failure modes above the script cannot detect

0

Of them produce a non-zero exit code

1

Test that finds most of them

Point it at a replica, and then watch the replica

One change is worth making before any of the above, and it is free if you already have a read replica: dump the replica, not the primary.

A dump is a long read holding a snapshot open for its whole duration. On the primary that means competing with production for I/O, blocking exclusive DDL, and holding back VACUUM so dead tuples pile up for however many hours it runs. On a small database nobody notices. On a large one it is the reason the backup window is a thing people schedule around.

Then it introduces a failure mode worse than the six above, and almost nobody instruments it.

A stalled replica backs up stale data, silently. Replication stops. Nothing alerts, because nothing is watching. Every nightly dump afterwards is a faithful, correctly-sized, perfectly restorable copy of a database frozen at the moment replication broke. You find out when you restore and the data is six weeks old.

The check that catches it

A stalled replica does not produce small backups, it produces identical ones. So the size comparison that catches a truncated dump will not catch this. You need the opposite alarm: a dump whose size stops changing at all. pg_last_xact_replay_timestamp() on the standby is the direct answer, and it takes one line to record alongside each run.

One more thing specific to Postgres. A long pg_dump against a hot standby can be killed mid-run by WAL replay with canceling statement due to conflict with recovery. hot_standby_feedback = on fixes it at the cost of bloat back on the primary you were trying to protect. max_standby_streaming_delay = -1 fixes it by letting the standby fall behind while the dump runs, which is usually the right trade for a node whose job is backups.

What fixing each one actually requires

The uncomfortable part is that every fix is small, and there are a lot of them.

FailureThe fix
Useless dumpRecord size and duration per run, alert on deviation
Stale replicaRecord replay lag per run, alert when the size stops moving
SilencePush a heartbeat outward, alert on the absence of a run
OverlapA scheduler that knows whether the previous run finished
RetentionExpiry that is aware of whether newer copies exist
CredentialsA secret store and short-lived credentials, not an env file
Blast radiusA destination in a different account with different credentials

None of these is hard. Together they are a system that somebody on your team now owns, and it is nobody's favourite work.

What we do instead

We run the same dump. The difference is what surrounds it.

Every run produces a record: what ran, how long it took, how many bytes it produced, where each copy landed, and what failed if anything did. Runs are durable, so a worker dying halfway through does not quietly skip a night. Retention is expressed as a policy rather than a lifecycle rule, so expiry knows about your other copies. Credentials live in a secret store and are read at run time by the worker that needs them.

The bytes go to a bucket you own, compressed and encrypted before they leave the machine that produced them.

Do this before you evaluate anything

Take last night's dump from whatever you have now, restore it into a scratch database, and compare row counts on your three biggest tables. Most teams discover something in that hour. Whatever you find will tell you more about your real exposure than any vendor comparison.

The honest summary

If you have one database and one engineer who remembers to check, the script above is genuinely fine. The moment you have five sources, two clouds, or an auditor asking when the last successful restore was, you are building the second column of that table whether you meant to or not.

Read next

How to back up the output of any script

Every backup tool has a list of things it supports, and your thing is not on it. The honest answer is an escape hatch. If you can write a script that produces a file, that file can be backed up on a schedule like anything else.

Backups that survive the thing that took out production.

How it worksStart free