Compression
Per-backup compression, applied before encryption and often inside it.
View as MarkdownCompression is off by default and set per backup.
backups:
- name: prod-db
compression: trueIn a manifest it is a switch, not a block. apply sends gzip at level 6, which is the
default everywhere. Use the CLI when you want a different level:
sctl backup configure <backup-id> --compression --compression-level 6On a worker's own config.yaml it is a switch too, one per backup:
backups:
"<backup-uuid>":
compression: trueWhat actually runs
There is one algorithm: gzip, or its close relative zlib when the artifact is also
encrypted. algo is accepted and stored so a manifest can be explicit, but gzip is what
runs regardless of what you write there.
level | Behaviour |
|---|---|
1 | Fastest, largest output |
6 | A good default for database dumps |
9 | Smallest, and disproportionately slower |
0 or unset | The library default |
Anything outside 1 to 9 falls back to the default rather than being refused.
A local worker ignores level and compresses at gzip's default. The setting is carried
on the definition and honoured by our own pipeline; the worker's config has a switch and
nothing else. If a specific level matters on a local backup, produce the artifact yourself
through a script source.
Order matters
Compress, then encrypt. Never the other way round.
Ciphertext is indistinguishable from random data, and random data does not compress. A 2 GB dump encrypted first and then compressed is 2 GB plus a gzip header. The pipeline enforces the order, so this is not something you can get wrong, but it is worth knowing why the order is not configurable.
inspect → compress → encrypt → deliver → finalizeEach layer that ran adds its own suffix, in the order it was applied, so the name tells you how to peel it.
| Backup settings | Filename ends | Artifact reports |
|---|---|---|
| Neither | app.dump | compressed: false, encrypted: false |
| Compression only | app.dump.gz | compressed: true |
| Encryption only | app.dump.gpg | encrypted: true |
| Both | app.dump.gz.gpg | compressed: true, encrypted: true |
The fourth row is the one to remember when recovering by hand: gpg --decrypt leaves you
holding the .gz, and there is a gunzip still to run. See
Artifact format.
Encryption compresses on its own
A PGP message compresses its payload as part of being built. That happens whenever the worker
has an encryption key, independently of the compression setting, and gpg undoes it
transparently on the way out. It is not the layer named in the filename and it is not
something you ever act on.
On a local backup that the worker encrypts, turning compression on buys little. The
gzip pass runs first and the PGP layer then compresses its output again, which is a second
pass over data that is already small. Leave it off unless you have measured a case where it
helps, such as a folder of text that gzip handles far better than zlib.
The practical reading: if you run local backups with encryption, you are already getting compression, and the transfer out of your site is already smaller.
Is it worth it
Almost always, for text-shaped data.
| Source | Typical reduction |
|---|---|
pg_dump -Fc | Little. The custom format already compresses |
mysqldump SQL | Large. It is plain text |
folder of documents or code | Large |
folder of images, video or archives | None, and you pay CPU for it |
redis RDB | Moderate |
s3 tar of already-compressed objects | None |
Two cases where turning it off is right:
- The source is already compressed. JPEGs, MP4s, parquet, or a
pg_dumpcustom-format file. You spend worker CPU to save nothing. - The dump is enormous and the window is tight. Level 9 on a 500 GB dump can cost more wall-clock time than the storage it saves is worth.
The trade
| You pay | You save |
|---|---|
| CPU time on whichever machine runs the step | archive: fewer bytes held, every day, for the artifact's whole life |
| Wall-clock time inside the run's timeout | transit: fewer bytes egressed per delivered copy |
storage: fewer bytes while the run works |
Compression is billed for once and saves every day the artifact exists. On a 90-day retention that is a good trade for anything text-shaped.
Compression happens inside the pipeline's step timeouts. If a run starts failing at the compress step after you raise the level, lower it rather than assuming the source got bigger.
Changing it
Compression is not write-once. Change it whenever you like; it applies to future runs and never to artifacts already stored.
An artifact carries its own compressed flag, so a backup whose setting changed halfway
through its life produces a mix, and each artifact still says what it is. Nothing about a
restore needs you to remember which setting was in force.