Artifact format
Self-describing, with no proprietary layer to reverse.
View as MarkdownThere is no saved.sh format. An artifact is a dump wrapped in at most two standard layers, both of which predate us by decades and neither of which needs anything of ours to unwrap.
This page tells you how to identify what you are holding and peel it, with no account.
The layers
[ PGP message [ compressed [ the dump ] ] ]| Layer | What it is | Unwrapped by |
|---|---|---|
| Encryption | OpenPGP (RFC 4880), public-key | gpg --decrypt |
| Compression | gzip (RFC 1952) | gunzip |
| The dump | Whatever the source type produces | The source's own tool |
Both wrappers are optional and set per backup. Compression is applied first, so an artifact with both is a gzip file inside a PGP message, and you unwrap in that order.
The four combinations
This is the table that matters, because the compression layer moves depending on whether encryption is on.
| Backup settings | Filename ends | Structure | To open |
|---|---|---|---|
| Neither | app.dump | The dump | Nothing to do |
| Compression only | app.dump.gz | gzip(dump) | gunzip |
| Encryption only | app.dump.gpg | pgp(dump) | gpg --decrypt |
| Both | app.dump.gz.gpg | pgp(gzip(dump)) | gpg --decrypt, then gunzip |
gpg --decrypt on a compressed artifact hands you a gzip file, not the dump. The name
says so: every suffix is one layer, in the order it was applied, so app.dump.gz.gpg peels
right to left. Feeding that gzip stream to pg_restore or mysql produces a confusing
parse error, not a useful one.
gpg --decrypt app.dump.gz.gpg > app.dump.gz
gunzip app.dump.gzPGP compresses its own payload as well, and gpg undoes that transparently. That inner
layer is not the gzip above it, and it is not something you act on.
Filenames
The name is built by appending to whatever the source produced.
| Source | Base name |
|---|---|
postgres | <database>.dump |
mysql | <database>.sql |
redis | <host>.redis.jsonl, or redis.redis.jsonl |
s3 | <bucket>.tar |
folder | <folder>.zip |
file | The file's own name |
web | From Content-Disposition, or the URL's last path segment |
script | Your script's own name, without its extension |
Names are sanitised to letters, digits, -, _ and ., and capped in length. A name is a
convenience, not a contract: the metadata is the contract, and the filename can be
absent entirely without affecting anything.
sctl artifact get <artifact-id>Filename: app.dump.gz.gpg
Size: 1483920128 bytes
Original: 4192104448 bytes
Checksum: sha256:9f86d081884c7d65...
Compressed: yes
Encrypted: yes
Key: 3AA5C34371567BD2Identifying an artifact from its bytes
If all you have is the file, the first few bytes tell you what it is. This works with no account, no metadata and no network.
# Linux and macOS
file artifact # often enough on its own
xxd -l 16 artifact # the definitive answer# Windows PowerShell: there is no `file`, so read the bytes
Format-Hex .\artifact -Count 16Peeling in the wrong order is the usual mistake: a PGP message that decrypts to something
file calls gzip compressed data is correct and expected, not a sign you used the wrong key.
| First bytes | Hex | It is |
|---|---|---|
-----BEGIN PGP | An armoured PGP message. Text | |
| (binary) | 85 or 84, or C1 | A binary PGP message |
| (binary) | 1F 8B | gzip |
PGDMP | 50 47 44 4D 50 | A Postgres custom-format dump |
PK\x03\x04 | 50 4B 03 04 | A zip, so a folder artifact |
REDIS | 52 45 44 49 53 | A Redis RDB snapshot |
-- or /* | Plain SQL, so a mysql dump | |
| (at offset 257) | 75 73 74 61 72 | tar, so s3 |
For tar, the magic is not at the start:
# Linux and macOS
xxd -s 257 -l 5 artifact # expect "ustar"# Windows PowerShell
Format-Hex .\artifact -Offset 257 -Count 5Peel one layer, then look again. Two checks get you to the dump from any artifact we produce.
gpg --decrypt artifact > layer1 # if PGP
file layer1 # now what is it?Which key opens it
sctl artifact get reports Key, the fingerprint of the public key the artifact was
encrypted to. So does the file itself, without us:
gpg --list-packets artifact | head -3:pubkey enc packet: version 3, algo 1, keyid 3AA5C34371567BD2That key ID is what you match against your private keys. It is the field that matters after a key rotation, when different artifacts of the same backup need different keys.
Verifying the checksum
The recorded checksum is SHA-256 of the file as it was uploaded, which is not always the file you download.
| Backup | Checksum matches |
|---|---|
local, encrypted on the worker | The downloaded file, directly |
cloud or manual, with our compression or encryption | The file after you decrypt and decompress |
| Anything with no transforms at all | The downloaded file |
# Linux
sha256sum artifact
# macOS
shasum -a 256 artifact# Windows PowerShell
Get-FileHash .\artifact -Algorithm SHA256For a cloud or manual backup with transforms, decrypt first and hash the result:
gpg --decrypt artifact > restored
sha256sum restored # shasum -a 256 on macOSThe reason is that the hash is taken when the bytes arrive, before we transform them, and it is deliberately never recomputed afterwards: it attests to what you gave us, which is the thing worth attesting to. A mismatch at upload time fails the run outright, so an artifact that exists has already been verified once.
sctl artifact get also reports Original, the uploaded size, alongside Size, the stored
size. When they differ, our pipeline transformed the file and the second row of the table
above applies.
Object layout
Delivered to a bucket you own, an artifact lands at a predictable key:
<prefix>/<workspace-id>/<backup-id>/<run-id>/artifactThe layout matches our own, so the same tree reads the same way in either bucket. There is one object per run, with no sidecar, manifest or index file. Everything you need to interpret it is in the bytes and in the run ID embedded in its path.
aws s3 ls s3://my-bucket/<workspace-id>/<backup-id>/ --recursive
aws s3 cp s3://my-bucket/<workspace-id>/<backup-id>/<run-id>/artifact ./artifactWhy there is no container format
A container would let us record the source type, the layer order and the schema version in one place, and it would be more convenient about eleven months out of twelve.
It would also be a format only our tooling reads. In the twelfth month, when the thing that went wrong is us, you would need a working saved.sh to parse your own backup. That is precisely the failure the product exists to prevent, so the convenience is not worth it.
The consequence is that identifying a bare artifact takes a file command instead of a
header read. That is the trade, and it is deliberate.