Three words get used as though they mean roughly the same thing, and they get budgeted as though buying one covers the others. They do not, and the confusion tends to surface at the worst possible moment.
The shortest version: replication protects you from hardware, snapshots protect you from yourself, and backups protect you from everything else.
The distinction that matters
| Replication | Snapshot | Backup | |
|---|---|---|---|
| Copies your mistake | Yes, in milliseconds | No | No |
| Lives outside the system | No | No | Yes |
| Survives account loss | No | No | Yes |
| Restorable to a different provider | No | Rarely | Yes |
| Cost | Roughly 1x your primary | Cheap, incremental | Cheap, and compresses well |
| Recovery speed | Instant failover | Minutes | Minutes to hours |
Read the first row carefully, because it is the one that surprises people.
Replication is not a time machine
Replication exists to keep a second copy identical to the first. That is its entire job, and it is very good at it.
Which means when you run DROP TABLE users on the primary, the replica applies
DROP TABLE users too, correctly, immediately, and with no way to decline. A
replica is not a copy of your data as it was. It is a copy of your data as it is,
including the part that just went wrong.
Replication answers exactly one question: what if this machine dies. That is a real question and a real answer. It is not the question you are asking when you say "backup".
Snapshots are an undo button with a boundary problem
A snapshot is a point-in-time copy, usually incremental, usually stored by the same system that stores your data.
They are genuinely excellent for the most common incident, which is somebody deleting something at 11am and noticing at 11:05. Restore from this morning, lose five minutes, move on.
The limitation is structural rather than technical. Snapshots live inside the account, inside the provider, in the provider's format. Every failure that takes the boundary takes the snapshots with it:
- The account is suspended for a billing problem.
- Someone with admin credentials deletes the instance, and the snapshots go too.
- The provider has a bad day in your region.
- You want to leave the provider, and discover the snapshots cannot come.
1
Question replication answers
1
Question snapshots answer
4
Ways the boundary takes both at once
What makes something a backup
Three properties, and a copy needs all three:
- It is separate. Different account, different credentials, ideally a different provider. If one compromise reaches both, you have one copy.
- It is a point in time you chose. Not "whatever the primary looks like now", but a state you can name and return to.
- It is openable without the thing that made it. An open format, restorable with standard tools, by someone who has never heard of your vendor.
The third one gets skipped most often and hurts most. A copy you can only read through the vendor that made it is a copy whose availability is the vendor's uptime and the vendor's goodwill.
Use all three, for different jobs
This is not a competition, and the right answer for most teams is all of them:
- Replication for availability. Node dies, traffic moves, nobody notices.
- Snapshots for fast recovery from ordinary mistakes, which are the majority of incidents by count.
- Backups for everything that reaches past the boundary, which are the minority of incidents by count and the majority by consequence.
Where teams get into trouble is buying the first two and writing "backups" on the line item.
Our line on it
We do the third one, deliberately and only. Backups land in a bucket you own, in an open format, encrypted before they leave the machine that made them, with a documented path to open them using standard tools and no software of ours.
Keep your replicas. Keep your snapshots. They are doing jobs we are not trying to do.