---
title: "Data flow"
description: "Bulk data moves directly to object storage. Only metadata passes through us."
url: "https://saved.sh/docs/security/data-flow"
---

Your data and our orchestration travel on different paths, and they meet only at an object
store neither side reads through the other.

This is a structural decision rather than a policy one: the bulk path physically does not
route through our API, so there is no code path we could add later that would read it.

## The local path [#the-local-path]

```
your host                              us
─────────                              ──
dump ──► encrypt ──► checksum ──► upload ──────────────► object storage
                                     │
                                     └── run id, checksum, size, filename ──► our API
```

| Leg                     | Carries                                               | Who sees it             |
| ----------------------- | ----------------------------------------------------- | ----------------------- |
| Dump, encrypt, checksum | Your data, in plaintext then ciphertext               | Your host only          |
| Upload                  | **Ciphertext**, via a presigned URL scoped to one run | Object storage          |
| Confirm                 | Run ID, checksum, byte count, filename                | Our API                 |
| Orchestration           | Workspace ID, backup ID, run ID                       | Our orchestration layer |

**The scheduling task carries three identifiers and nothing else.** No credentials, no source
configuration, no data are ever sent to your worker, because none of it came from us: the
worker resolves the backup ID against its own config file.

## The cloud path [#the-cloud-path]

```
                    us
                    ──
your source ──► dump ──► compress ──► encrypt ──► object storage
```

Everything happens on our infrastructure. We hold the source credential in our vault, read it
at run time, and handle plaintext until the encrypt step.

That is the trade a cloud backup makes, stated on the
[threat model](/docs/security/threat-model#the-trade-you-make-by-kind).

## What crosses which boundary [#what-crosses-which-boundary]

|                                    | Local                         | Cloud                  | Manual                     |
| ---------------------------------- | ----------------------------- | ---------------------- | -------------------------- |
| Source credential                  | Your disk only                | **Our vault**          | None                       |
| Plaintext                          | Your host only                | **Our infrastructure** | **Our infrastructure**     |
| Private key                        | Yours only                    | Yours only             | Yours only                 |
| Artifact bytes                     | Ciphertext, direct to storage | Direct to storage      | Direct from you to storage |
| Checksum, size, filename           | Our API                       | Our API                | Our API                    |
| Backup name, source type, schedule | Ours                          | Ours                   | Ours                       |

## Uploads and downloads never pass through our API [#uploads-and-downloads-never-pass-through-our-api]

Both directions use **presigned URLs**: a time-limited, single-object grant that the client
uses to talk to object storage directly.

| Direction    | How                                                                      |
| ------------ | ------------------------------------------------------------------------ |
| Upload       | `POST /v1/runs/{run}/upload-url`, then `PUT` straight to storage         |
| Large upload | Multipart, with part URLs issued in batches and completed by ETag        |
| Download     | `POST /v1/artifacts/{id}/download-url`, then `GET` straight from storage |

Two consequences worth stating:

* **The worker never holds long-lived storage credentials.** It gets a URL for one object,
  for one run.
* **We do not observe the transfer.** We know a URL was issued, which is why
  `artifact.download_url_issued` is the one read the [audit log](/docs/security/audit-log)
  records. The fetch itself is between you and the object store.

## What the orchestration layer sees [#what-the-orchestration-layer-sees]

Only control information moves through it, and "only control information" is still more than
nothing. Being precise about it:

| What          | Detail                                             |
| ------------- | -------------------------------------------------- |
| Identifiers   | Workspace, backup and run IDs                      |
| Progress      | Byte counts, object counts, upload part numbers    |
| Scratch paths | Progress messages name the local staging file      |
| Failure text  | Up to 2000 characters of the failing tool's stderr |

<Callout type="warn">
  That last row is the one to think about. When `pg_dump` fails, we keep its own message,
  because "access denied for user [backup@db.internal](mailto:backup@db.internal)" is worth vastly more to you than "exit
  status 2". It can contain hostnames, usernames, and database or table names. It does not
  contain the password, which is passed through the environment and never on a command line.
</Callout>

If that trade is wrong for you, the lever is a [`script` source](/docs/backups/sources/script):
handle errors yourself and print only what you want recorded.

## Tenancy [#tenancy]

Every credential is scoped to one workspace, and the boundary is enforced at each layer
rather than by convention.

| Layer         | Boundary                                                                        |
| ------------- | ------------------------------------------------------------------------------- |
| API           | Every workspace-scoped route checks a permission on the caller's workspace      |
| Orchestration | One namespace per workspace; a worker key is bound to its own                   |
| Task routing  | One queue per worker, so a run reaches only the machine holding its credentials |
| Object keys   | `<workspace-id>/<backup-id>/<run-id>/artifact`                                  |
| Audit         | `workspace_id` is the query key on every row                                    |

A worker's connection may poll its own queue, report results, and describe its own namespace.
It cannot start work, administer anything, or see another queue, let alone another workspace.

## The direct-read question [#the-direct-read-question]

Artifacts in **our** storage are reachable only through a presigned URL. We do not hand out
bucket credentials, which is the correct posture and is also a dependency on us.

If your recovery plan must survive us being unreachable, add a
[destination you own](/docs/backups/delivery). The artifact lands in your bucket under the
same key layout, readable with your own credentials and any S3 client:

```
<prefix>/<workspace-id>/<backup-id>/<run-id>/artifact
```

Same finished object, compressed and encrypted exactly as ours is.

## Next [#next]

<Cards>
  <Card href="/docs/workers/local-backups" title="What stays local" description="The same boundary, from the worker's side." />

  <Card href="/docs/security/threat-model" title="Threat model" description="What these boundaries do and do not stop." />

  <Card href="/docs/backups/delivery" title="Delivery" description="Keeping a copy you can reach without us." />
</Cards>
