---
title: "Artifact format"
description: "Self-describing, with no proprietary layer to reverse."
url: "https://saved.sh/docs/recover/artifact-format"
---

There is no saved.sh format. An artifact is a dump wrapped in at most two standard layers,
both of which predate us by decades and neither of which needs anything of ours to unwrap.

This page tells you how to identify what you are holding and peel it, with no account.

## The layers [#the-layers]

```
[ PGP message  [ compressed  [ the dump ] ] ]
```

| Layer       | What it is                        | Unwrapped by          |
| ----------- | --------------------------------- | --------------------- |
| Encryption  | OpenPGP (RFC 4880), public-key    | `gpg --decrypt`       |
| Compression | gzip (RFC 1952)                   | `gunzip`              |
| The dump    | Whatever the source type produces | The source's own tool |

Both wrappers are optional and set per backup. **Compression is applied first**, so an
artifact with both is a gzip file inside a PGP message, and you unwrap in that order.

## The four combinations [#the-four-combinations]

This is the table that matters, because the compression layer moves depending on whether
encryption is on.

| Backup settings  | Filename ends     | Structure       | To open                        |
| ---------------- | ----------------- | --------------- | ------------------------------ |
| Neither          | `app.dump`        | The dump        | Nothing to do                  |
| Compression only | `app.dump.gz`     | gzip(dump)      | `gunzip`                       |
| Encryption only  | `app.dump.gpg`    | pgp(dump)       | `gpg --decrypt`                |
| **Both**         | `app.dump.gz.gpg` | pgp(gzip(dump)) | `gpg --decrypt`, then `gunzip` |

<Callout type="warn">
  **`gpg --decrypt` on a compressed artifact hands you a gzip file, not the dump.** The name
  says so: every suffix is one layer, in the order it was applied, so `app.dump.gz.gpg` peels
  right to left. Feeding that gzip stream to `pg_restore` or `mysql` produces a confusing
  parse error, not a useful one.

  ```bash
  gpg --decrypt app.dump.gz.gpg > app.dump.gz
  gunzip app.dump.gz
  ```

  PGP compresses its own payload as well, and `gpg` undoes that transparently. That inner
  layer is not the gzip above it, and it is not something you act on.
</Callout>

## Filenames [#filenames]

The name is built by appending to whatever the source produced.

| Source     | Base name                                                  |
| ---------- | ---------------------------------------------------------- |
| `postgres` | `<database>.dump`                                          |
| `mysql`    | `<database>.sql`                                           |
| `redis`    | `<host>.redis.jsonl`, or `redis.redis.jsonl`               |
| `s3`       | `<bucket>.tar`                                             |
| `folder`   | `<folder>.zip`                                             |
| `file`     | The file's own name                                        |
| `web`      | From `Content-Disposition`, or the URL's last path segment |
| `script`   | Your script's own name, without its extension              |

Names are sanitised to letters, digits, `-`, `_` and `.`, and capped in length. A name is a
convenience, not a contract: **the metadata is the contract**, and the filename can be
absent entirely without affecting anything.

```bash
sctl artifact get <artifact-id>
```

```
Filename:    app.dump.gz.gpg
Size:        1483920128 bytes
Original:    4192104448 bytes
Checksum:    sha256:9f86d081884c7d65...
Compressed:  yes
Encrypted:   yes
Key:         3AA5C34371567BD2
```

## Identifying an artifact from its bytes [#identifying-an-artifact-from-its-bytes]

If all you have is the file, the first few bytes tell you what it is. This works with no
account, no metadata and no network.

```bash
# Linux and macOS
file artifact          # often enough on its own
xxd -l 16 artifact     # the definitive answer
```

```powershell
# Windows PowerShell: there is no `file`, so read the bytes
Format-Hex .\artifact -Count 16
```

Peeling in the wrong order is the usual mistake: a PGP message that decrypts to something
`file` calls `gzip compressed data` is correct and expected, not a sign you used the wrong key.

| First bytes      | Hex                   | It is                             |
| ---------------- | --------------------- | --------------------------------- |
| `-----BEGIN PGP` |                       | An **armoured** PGP message. Text |
| (binary)         | `85` or `84`, or `C1` | A **binary** PGP message          |
| (binary)         | `1F 8B`               | gzip                              |
| `PGDMP`          | `50 47 44 4D 50`      | A Postgres custom-format dump     |
| `PK\x03\x04`     | `50 4B 03 04`         | A zip, so a `folder` artifact     |
| `REDIS`          | `52 45 44 49 53`      | A Redis RDB snapshot              |
| `--` or `/*`     |                       | Plain SQL, so a `mysql` dump      |
| (at offset 257)  | `75 73 74 61 72`      | tar, so `s3`                      |

For tar, the magic is not at the start:

```bash
# Linux and macOS
xxd -s 257 -l 5 artifact    # expect "ustar"
```

```powershell
# Windows PowerShell
Format-Hex .\artifact -Offset 257 -Count 5
```

Peel one layer, then look again. Two checks get you to the dump from any artifact we
produce.

```bash
gpg --decrypt artifact > layer1     # if PGP
file layer1                         # now what is it?
```

## Which key opens it [#which-key-opens-it]

`sctl artifact get` reports `Key`, the fingerprint of the public key the artifact was
encrypted to. So does the file itself, without us:

```bash
gpg --list-packets artifact | head -3
```

```
:pubkey enc packet: version 3, algo 1, keyid 3AA5C34371567BD2
```

That key ID is what you match against your private keys. It is the field that matters after
a key rotation, when different artifacts of the same backup need different keys.

## Verifying the checksum [#verifying-the-checksum]

The recorded checksum is SHA-256 of the file **as it was uploaded**, which is not always the
file you download.

| Backup                                                  | Checksum matches                              |
| ------------------------------------------------------- | --------------------------------------------- |
| `local`, encrypted on the worker                        | The **downloaded** file, directly             |
| `cloud` or `manual`, with our compression or encryption | The file **after** you decrypt and decompress |
| Anything with no transforms at all                      | The downloaded file                           |

```bash
# Linux
sha256sum artifact

# macOS
shasum -a 256 artifact
```

```powershell
# Windows PowerShell
Get-FileHash .\artifact -Algorithm SHA256
```

For a cloud or manual backup with transforms, decrypt first and hash the result:

```bash
gpg --decrypt artifact > restored
sha256sum restored          # shasum -a 256 on macOS
```

The reason is that the hash is taken when the bytes arrive, before we transform them, and it
is deliberately never recomputed afterwards: it attests to **what you gave us**, which is the
thing worth attesting to. A mismatch at upload time fails the run outright, so an artifact
that exists has already been verified once.

`sctl artifact get` also reports `Original`, the uploaded size, alongside `Size`, the stored
size. When they differ, our pipeline transformed the file and the second row of the table
above applies.

## Object layout [#object-layout]

Delivered to a bucket you own, an artifact lands at a predictable key:

```
<prefix>/<workspace-id>/<backup-id>/<run-id>/artifact
```

The layout matches our own, so the same tree reads the same way in either bucket. There is
one object per run, with no sidecar, manifest or index file. Everything you need to
interpret it is in the bytes and in the run ID embedded in its path.

```bash
aws s3 ls s3://my-bucket/<workspace-id>/<backup-id>/ --recursive
aws s3 cp s3://my-bucket/<workspace-id>/<backup-id>/<run-id>/artifact ./artifact
```

## Why there is no container format [#why-there-is-no-container-format]

A container would let us record the source type, the layer order and the schema version in
one place, and it would be more convenient about eleven months out of twelve.

It would also be a format only our tooling reads. In the twelfth month, when the thing that
went wrong is us, you would need a working saved.sh to parse your own backup. That is
precisely the failure the product exists to prevent, so the convenience is not worth it.

The consequence is that identifying a bare artifact takes a `file` command instead of a
header read. That is the trade, and it is deliberate.
