Threat model
The attacker this design assumes, and the ones it does not stop.
View as MarkdownThe attacker we design against is whoever holds your production access. Not a stranger on the internet: the credential that is already inside, whether it belongs to an intruder, a departing employee, or an automation with more permission than anyone realised.
That choice explains the whole architecture. Defending against an outsider is a firewall problem. Defending against your own credentials is a fate-sharing problem, and the answer is to put the copy somewhere those credentials do not reach.
What we defend against
| Threat | Defence | Where |
|---|---|---|
| Production compromised, attacker deletes the backups | The archive is not reachable from production, and holds different credentials | Data flow |
| Attacker with your cloud account destroys snapshots | Copies live outside that account, under our keys or a destination you chose | Delivery |
| We are compromised, or an employee of ours is curious | Local backups are encrypted before upload; we hold ciphertext and no key | Encryption |
| A leaked worker key is used to attack the workspace | The key carries two permissions and is bound to one namespace and one queue | Permissions |
| A leaked API key mints further credentials | api-keys:*, workers:* and billing:* are refused on machine keys, at two layers | API keys |
| Someone deletes an artifact to cover their tracks | Deletion is refused while locked, and the attempt is on the audit trail | Audit log |
| Ransomware encrypts the primary and the snapshots | The artifact is already elsewhere, and retention is write-once | Retention |
| A tenant reads another tenant's data | Every credential is organization-scoped; the namespace is the workspace | Authentication |
What we do not defend against
This is the more useful list, and it is deliberately not buried.
Your worker's host
A local backup runs on your machine, reads your sources with your credentials, and holds plaintext on your disk between the dump and the encrypt step. Anyone with root on that host has your data, and nothing in our design changes that.
Run the worker somewhere you would already be comfortable holding a database dump, because for a few minutes per run it does.
Your private key, and anyone who has it
The key is the whole of the confidentiality guarantee. Whoever holds the private half can read every artifact it opens, forever, including artifacts they were never authorised to download. We cannot revoke a key we never had.
See key management.
An artifact already downloaded
Once the bytes are on someone's laptop, they are on someone's laptop. The audit log records that a download URL was issued, which is the last moment we can observe. Everything after that is outside our view by construction, because the fetch goes straight to object storage and never passes through us.
Storage-layer deletion by whoever holds the archive credentials
lock_for is enforced by our application: it stops our API, our dashboard, our support
tooling and our retention sweep. It is not object-lock enforcement at the bucket.
Read plainly: an attacker who obtained our archive storage credentials would not be stopped by an artifact lock. It is a control against mistake and misuse, not a WORM guarantee. If you need storage-layer immutability, deliver to a bucket you own and configure object lock there, where enforcement is yours.
A backup that was never encrypted
Encryption is optional and nothing warns you at run time. A backup configured without a key
produces plain artifacts indefinitely, and they sit in our storage as plaintext. Check
Encrypted on an artifact rather than assuming.
The correctness of your own dump
We back up what your source gives us. If the credential could only see half the tables, or
your script source wrote a log instead of data, the run succeeds and the artifact is
useless. That failure is invisible until a restore, which is why
rehearsal is a page and not a footnote.
Availability of your source
A backup does not make your database more available. It makes losing it survivable.
The trade you make by kind
The choice of local versus cloud is a threat-model decision, not a convenience one.
local | cloud | |
|---|---|---|
| Who holds the source credential | You | Us, in our vault |
| Who sees plaintext | Your host | Our infrastructure |
| Compromise of us exposes | Ciphertext and metadata | The source credential and the plaintext |
| Compromise of your worker exposes | The source credential and plaintext | Nothing of ours |
| Requires you to operate a process | Yes | No |
Cloud backups are a real transfer of trust, and we would rather say so than imply the two are equivalent. They exist because sometimes you would rather we connected than stand up a process, and that is a legitimate choice as long as it is made deliberately.
If the source is sensitive and you can run a worker, run a worker.
What we can see, in either case
| We always see | We never see (local) | We see (cloud) |
|---|---|---|
| Backup names, source types, schedules | Source credentials | Source credentials, in our vault |
| Artifact sizes, checksums, timestamps | Plaintext | Plaintext, during the run |
| Run history, step names, failure text | Your private key | |
| Which key ID an artifact used |
Names and sizes are metadata we necessarily hold. If a database name is itself sensitive, that is worth knowing before you schedule it.
Assumptions
The model rests on four things being true. Each is worth checking against your own situation.
- You keep the private key somewhere that does not share fate with production. A key on the database server defeats the design.
- You run at least one backup whose copies you can reach without us. Otherwise our availability is in your recovery path.
- You rehearse. An unrestored backup is a hypothesis, and the model assumes you have turned it into a fact.
- Your worker host is trusted. Everything on the local path depends on it.