Data flow
Bulk data moves directly to object storage. Only metadata passes through us.
View as MarkdownYour data and our orchestration travel on different paths, and they meet only at an object store neither side reads through the other.
This is a structural decision rather than a policy one: the bulk path physically does not route through our API, so there is no code path we could add later that would read it.
The local path
your host us
───────── ──
dump ──► encrypt ──► checksum ──► upload ──────────────► object storage
│
└── run id, checksum, size, filename ──► our API| Leg | Carries | Who sees it |
|---|---|---|
| Dump, encrypt, checksum | Your data, in plaintext then ciphertext | Your host only |
| Upload | Ciphertext, via a presigned URL scoped to one run | Object storage |
| Confirm | Run ID, checksum, byte count, filename | Our API |
| Orchestration | Workspace ID, backup ID, run ID | Our orchestration layer |
The scheduling task carries three identifiers and nothing else. No credentials, no source configuration, no data are ever sent to your worker, because none of it came from us: the worker resolves the backup ID against its own config file.
The cloud path
us
──
your source ──► dump ──► compress ──► encrypt ──► object storageEverything happens on our infrastructure. We hold the source credential in our vault, read it at run time, and handle plaintext until the encrypt step.
That is the trade a cloud backup makes, stated on the threat model.
What crosses which boundary
| Local | Cloud | Manual | |
|---|---|---|---|
| Source credential | Your disk only | Our vault | None |
| Plaintext | Your host only | Our infrastructure | Our infrastructure |
| Private key | Yours only | Yours only | Yours only |
| Artifact bytes | Ciphertext, direct to storage | Direct to storage | Direct from you to storage |
| Checksum, size, filename | Our API | Our API | Our API |
| Backup name, source type, schedule | Ours | Ours | Ours |
Uploads and downloads never pass through our API
Both directions use presigned URLs: a time-limited, single-object grant that the client uses to talk to object storage directly.
| Direction | How |
|---|---|
| Upload | POST /v1/runs/{run}/upload-url, then PUT straight to storage |
| Large upload | Multipart, with part URLs issued in batches and completed by ETag |
| Download | POST /v1/artifacts/{id}/download-url, then GET straight from storage |
Two consequences worth stating:
- The worker never holds long-lived storage credentials. It gets a URL for one object, for one run.
- We do not observe the transfer. We know a URL was issued, which is why
artifact.download_url_issuedis the one read the audit log records. The fetch itself is between you and the object store.
What the orchestration layer sees
Only control information moves through it, and "only control information" is still more than nothing. Being precise about it:
| What | Detail |
|---|---|
| Identifiers | Workspace, backup and run IDs |
| Progress | Byte counts, object counts, upload part numbers |
| Scratch paths | Progress messages name the local staging file |
| Failure text | Up to 2000 characters of the failing tool's stderr |
That last row is the one to think about. When pg_dump fails, we keep its own message,
because "access denied for user backup@db.internal" is worth vastly more to you than "exit
status 2". It can contain hostnames, usernames, and database or table names. It does not
contain the password, which is passed through the environment and never on a command line.
If that trade is wrong for you, the lever is a script source:
handle errors yourself and print only what you want recorded.
Tenancy
Every credential is scoped to one workspace, and the boundary is enforced at each layer rather than by convention.
| Layer | Boundary |
|---|---|
| API | Every workspace-scoped route checks a permission on the caller's workspace |
| Orchestration | One namespace per workspace; a worker key is bound to its own |
| Task routing | One queue per worker, so a run reaches only the machine holding its credentials |
| Object keys | <workspace-id>/<backup-id>/<run-id>/artifact |
| Audit | workspace_id is the query key on every row |
A worker's connection may poll its own queue, report results, and describe its own namespace. It cannot start work, administer anything, or see another queue, let alone another workspace.
The direct-read question
Artifacts in our storage are reachable only through a presigned URL. We do not hand out bucket credentials, which is the correct posture and is also a dependency on us.
If your recovery plan must survive us being unreachable, add a destination you own. The artifact lands in your bucket under the same key layout, readable with your own credentials and any S3 client:
<prefix>/<workspace-id>/<backup-id>/<run-id>/artifactSame finished object, compressed and encrypted exactly as ours is.