---
title: "Compression"
description: "Per-backup compression, applied before encryption and often inside it."
url: "https://saved.sh/docs/backups/compression"
---

Compression is **off by default** and set per backup.

```yaml title="saved.yaml"
backups:
  - name: prod-db
    compression: true
```

**In a manifest it is a switch, not a block.** `apply` sends gzip at level 6, which is the
default everywhere. Use the CLI when you want a different level:

```bash
sctl backup configure <backup-id> --compression --compression-level 6
```

On a worker's own `config.yaml` it is a switch too, one per backup:

```yaml title="config.yaml"
backups:
  "<backup-uuid>":
    compression: true
```

## What actually runs [#what-actually-runs]

There is one algorithm: **gzip**, or its close relative zlib when the artifact is also
encrypted. `algo` is accepted and stored so a manifest can be explicit, but `gzip` is what
runs regardless of what you write there.

| `level`      | Behaviour                               |
| ------------ | --------------------------------------- |
| `1`          | Fastest, largest output                 |
| `6`          | A good default for database dumps       |
| `9`          | Smallest, and disproportionately slower |
| `0` or unset | The library default                     |

Anything outside 1 to 9 falls back to the default rather than being refused.

<Callout>
  **A local worker ignores `level` and compresses at gzip's default.** The setting is carried
  on the definition and honoured by our own pipeline; the worker's config has a switch and
  nothing else. If a specific level matters on a local backup, produce the artifact yourself
  through a [`script` source](/docs/backups/sources/script).
</Callout>

## Order matters [#order-matters]

**Compress, then encrypt.** Never the other way round.

Ciphertext is indistinguishable from random data, and random data does not compress. A
2 GB dump encrypted first and then compressed is 2 GB plus a gzip header. The pipeline
enforces the order, so this is not something you can get wrong, but it is worth knowing why
the order is not configurable.

```
inspect → compress → encrypt → deliver → finalize
```

Each layer that ran adds its own suffix, in the order it was applied, so the name tells you
how to peel it.

| Backup settings  | Filename ends     | Artifact reports                        |
| ---------------- | ----------------- | --------------------------------------- |
| Neither          | `app.dump`        | `compressed: false`, `encrypted: false` |
| Compression only | `app.dump.gz`     | `compressed: true`                      |
| Encryption only  | `app.dump.gpg`    | `encrypted: true`                       |
| Both             | `app.dump.gz.gpg` | `compressed: true`, `encrypted: true`   |

The fourth row is the one to remember when recovering by hand: `gpg --decrypt` leaves you
holding the `.gz`, and there is a `gunzip` still to run. See
[Artifact format](/docs/recover/artifact-format).

## Encryption compresses on its own [#encryption-compresses-on-its-own]

A PGP message compresses its payload as part of being built. That happens whenever the worker
has an encryption key, **independently of the `compression` setting**, and `gpg` undoes it
transparently on the way out. It is not the layer named in the filename and it is not
something you ever act on.

<Callout type="warn">
  **On a local backup that the worker encrypts, turning `compression` on buys little.** The
  gzip pass runs first and the PGP layer then compresses its output again, which is a second
  pass over data that is already small. Leave it off unless you have measured a case where it
  helps, such as a `folder` of text that gzip handles far better than zlib.
</Callout>

The practical reading: if you run local backups with encryption, you are already getting
compression, and the transfer out of your site is already smaller.

## Is it worth it [#is-it-worth-it]

Almost always, for text-shaped data.

| Source                                 | Typical reduction                            |
| -------------------------------------- | -------------------------------------------- |
| `pg_dump -Fc`                          | Little. The custom format already compresses |
| `mysqldump` SQL                        | Large. It is plain text                      |
| `folder` of documents or code          | Large                                        |
| `folder` of images, video or archives  | None, and you pay CPU for it                 |
| `redis` RDB                            | Moderate                                     |
| `s3` tar of already-compressed objects | None                                         |

Two cases where turning it off is right:

* **The source is already compressed.** JPEGs, MP4s, parquet, or a `pg_dump` custom-format
  file. You spend worker CPU to save nothing.
* **The dump is enormous and the window is tight.** Level 9 on a 500 GB dump can cost more
  wall-clock time than the storage it saves is worth.

## The trade [#the-trade]

| You pay                                     | You save                                                              |
| ------------------------------------------- | --------------------------------------------------------------------- |
| CPU time on whichever machine runs the step | `archive`: fewer bytes held, every day, for the artifact's whole life |
| Wall-clock time inside the run's timeout    | `transit`: fewer bytes egressed per delivered copy                    |
|                                             | `storage`: fewer bytes while the run works                            |

Compression is billed for once and saves every day the artifact exists. On a 90-day
retention that is a good trade for anything text-shaped.

<Callout type="warn">
  Compression happens inside the pipeline's step timeouts. If a run starts failing at the
  compress step after you raise the level, lower it rather than assuming the source got
  bigger.
</Callout>

## Changing it [#changing-it]

Compression is not write-once. Change it whenever you like; it applies to future runs and
never to artifacts already stored.

An artifact carries its own `compressed` flag, so a backup whose setting changed halfway
through its life produces a mix, and each artifact still says what it is. Nothing about a
restore needs you to remember which setting was in force.

## Next [#next]

<Cards>
  <Card href="/docs/backups/encryption" title="Encryption" description="The step compression folds into." />

  <Card href="/docs/recover/artifact-format" title="Artifact format" description="Unpacking each combination by hand." />

  <Card href="/docs/billing/usage" title="Usage" description="The meters compression moves." />
</Cards>
