---
title: "S3-compatible storage"
description: "Copy objects out of a bucket into one archive, local or cloud."
url: "https://saved.sh/docs/backups/sources/s3"
---

Lists a bucket or prefix and streams every object into a single **tar**. Available as both a
local and a cloud backup, and it needs no external tool: it speaks S3 directly.

Any S3-compatible store works: AWS S3, Cloudflare R2, MinIO, Backblaze B2, Wasabi,
DigitalOcean Spaces.

## What is taken [#what-is-taken]

Every object under the prefix, with its key as the tar entry name and its last-modified time
preserved.

| Included                                  | Not included                               |
| ----------------------------------------- | ------------------------------------------ |
| Object keys, contents, modification times | Object metadata and tags                   |
| The full prefix tree                      | Non-current versions                       |
|                                           | ACLs and bucket policy                     |
|                                           | Objects added after the listing pass began |

Keys ending in `/` are skipped: they are directory placeholders with no bytes.

The archive is **streamed** rather than staged, so backing up a 400 GB bucket does not need
400 GB of scratch space for the listing pass. Only the object being copied is in flight.

## Fields [#fields]

The same keys for both kinds. Local, in the worker's `config.yaml`:

| Key                 | Required | Notes                                                          |
| ------------------- | -------- | -------------------------------------------------------------- |
| `bucket`            | Yes      |                                                                |
| `access_key_id`     | Yes      |                                                                |
| `secret_access_key` | Yes      |                                                                |
| `region`            | No       | Defaults to `auto`, which is what R2 and several others want   |
| `endpoint`          | No       | For non-AWS stores. `https://` is added if you omit the scheme |
| `prefix`            | No       | Restrict to one path                                           |
| `use_path_style`    | No       | `true` for MinIO and most self-hosted stores                   |

Cloud, on the backup definition. Every field goes to the vault together and none is
returned by any endpoint:

| Fields                                                                                           |
| ------------------------------------------------------------------------------------------------ |
| `bucket`, `prefix`, `region`, `endpoint`, `use_path_style`, `access_key_id`, `secret_access_key` |

```yaml title="config.yaml"
backups:
  "<backup-uuid>":
    source_type: s3
    source:
      bucket: uploads
      endpoint: minio.internal:9000
      access_key_id: "<key>"
      secret_access_key: "<secret>"
      use_path_style: true
```

The credential needs `ListBucket` on the bucket and `GetObject` on the prefix. It never needs
write access, so do not grant any.

## Objects that change mid-copy [#objects-that-change-mid-copy]

A tar header declares the size before the bytes are written, and it cannot be corrected
afterwards. If an object is rewritten between the listing and the copy, its real size no
longer matches the header.

**The run fails rather than truncating the file.** The alternative would be an archive that
looks fine and contains a corrupt member, discovered during a restore.

```
uploads/report.pdf changed while being copied: listed 481920 bytes, read 502133
```

If this happens regularly, the prefix is actively written and a backup of it will never be
internally consistent. Back up a prefix that is append-only, or accept the failure rate and
retry.

An empty bucket or prefix fails with `EmptySource`, rather than producing an empty archive
you would mistake for a working backup.

## This is not delivery [#this-is-not-delivery]

Two different features look similar and are easy to confuse.

|                  | `s3` **source**               | A **destination**                                          |
| ---------------- | ----------------------------- | ---------------------------------------------------------- |
| Direction        | We read **from** your bucket  | We write **to** your bucket                                |
| Purpose          | Back the bucket's contents up | Store artifacts where you own them                         |
| Credential needs | `ListBucket`, `GetObject`     | `PutObject`, `GetObject`, and `DeleteObject` for the probe |
| Configured on    | The backup's source           | The workspace, then selected per backup                    |

You can use both at once: back up bucket A as a source, and deliver the resulting artifact to
bucket B as a destination. Do not point a backup's source and its destination at the same
bucket and prefix.

See [Delivery](/docs/backups/delivery).

## Sizing [#sizing]

The tar is the sum of the objects plus a small per-entry overhead. Compression depends
entirely on the contents: a bucket of JSON or CSV compresses well, a bucket of JPEGs and
MP4s does not compress at all and costs CPU to try.

Large buckets are the case where a local backup is worth preferring even when the store is
reachable from anywhere: the copy happens next to the data instead of across the internet.

## Restoring [#restoring]

```bash
sctl restore file <artifact-id> --path ./uploads.tar
tar -tvf ./uploads.tar          # inspect
tar -xf ./uploads.tar -C ./out  # extract
```

Then put the objects back with whatever tooling you normally use:

```bash
aws s3 sync ./out s3://uploads/
```

Restoring is deliberately not automatic. Writing several hundred thousand objects back into a
live bucket is not something a backup tool should do on your behalf.
