The database you care most about is frequently the one sitting on a private subnet with no route in from the internet. That is not an accident, it is the correct configuration, and it is exactly what makes backing it up awkward.
The usual options are all bad in the same way.
| Option | What it costs you |
|---|---|
| Give the database a public IP | You undo the decision you made deliberately |
| Open the port to a vendor's IP range | An allowlist that changes when the vendor's does, and one shared secret away from being a hole |
| Site-to-site VPN | Weeks of networking, and a permanent tunnel into your private network |
| Bastion with a scheduled SSH job | You are back to a cron job on a box, with the key material to prove it |
Every one of them treats the problem as "how do we get in".
Invert it
The pattern that avoids all four is the one your CI runners, your metrics agent and your log shipper already use: the thing inside makes an outbound connection and asks for work.
Nothing listens. No inbound rule, no allowlist, no tunnel, no public IP. The database keeps exactly the network position you gave it, and the only firewall change is the one you almost certainly already permit: outbound HTTPS on 443.
How it works here
You run a small worker process next to the database. On start it dials out and polls a queue that belongs only to it. When a run is due, the job comes down that connection, the worker dumps the database locally, and it uploads to your bucket.
Two properties of that arrangement are worth being specific about, because they are what make it safe rather than merely convenient.
Your credentials never leave the machine. The database password is read by the worker, on your side, from your configuration. It is not stored with us and it does not travel to us. Same for the bucket credentials.
Neither do the bytes. The steps of a run hand each other a path to a file on that machine. The dump does not pass through our orchestrator on its way to your bucket. Only metadata does: what ran, how long it took, how large the result was, and where it went.
0
Inbound firewall rules
0
Credentials that leave the machine
443
The only port involved
The one operational rule
A worker's queue is its identity, and only one process should poll it.
If you run two copies of the same worker configuration, the steps of a single run can split across two machines, and step two will look for a file that step one left somewhere else. That is not a subtle failure, it fails loudly, but it is worth knowing before you scale a deployment to two replicas out of habit.
Run one process per worker identity. If you want a second machine backing up different things, give it its own identity.
What this does not solve
Being honest about the boundary: an outbound worker does not help if the machine running it is itself the thing you are worried about. If an attacker owns that host, they own the database credentials and the worker process together.
That is the argument for the other half of the setup: send the copy to a bucket in a different account, and turn on Object Lock so that a credential holder cannot delete what is already written. Outbound polling solves the reachability problem. It was never going to solve the custody one.