---
title: "How to archive a website or a URL on a schedule"
description: "Downloading a URL is one curl away. Doing it every night, keeping the versions, noticing when the response silently becomes a login page, and storing it somewhere you control is the part that takes work."
url: "https://saved.sh/blog/archive-a-website-on-a-schedule"
date: "2026-08-08"
author: "saved.sh"
tag: "Engineering"
---

Plenty of things worth keeping are only available over HTTP. An export endpoint
on a SaaS you do not control. A published dataset that changes. A rendered report.
A partner's feed that they will regenerate and overwrite next Tuesday.

Fetching one is trivial. Everything after that is the work.

## Start with the curl that does not lie to you [#start-with-the-curl-that-does-not-lie-to-you]

```bash
curl --fail --location --silent --show-error \
  --retry 3 --retry-delay 5 \
  -H "Authorization: Bearer $TOKEN" \
  -o export.json \
  https://api.example.com/v1/export
```

`--fail` is the one people leave out, and it is the one that matters. Without it
curl exits zero on a 404 or a 500 and cheerfully writes the error body to your
file. Your archive fills up with plausibly-sized HTML pages that say "Sign in".

`--location` follows redirects, which is how an endpoint that moved keeps working.
`--retry` covers the transient 502 that would otherwise lose you a night.

## The four things that go wrong afterwards [#the-four-things-that-go-wrong-afterwards]

**The response quietly changes shape.** The endpoint starts returning a login
page, or an empty array, or a truncated body. All three produce a file. None of
them produce an error. The only defence is comparing this run's size against the
last one, which means something has to be recording the sizes.

**The token expires.** Usually on a weekend, usually silently, and usually
discovered when someone needs the data.

**Nobody keeps the versions.** `curl -o export.json` overwrites. If you want last
Tuesday's copy, you needed to have decided that last Tuesday.

**It lives on someone's laptop.** The most common home for a script like this is
a machine that gets replaced.

<Figure caption="The fetch is one command. The version history, the size comparison and the copy you control are the parts that take a system.">
  <ArtifactPipeline />
</Figure>

## What a web source does with it [#what-a-web-source-does-with-it]

We run essentially the curl above, on a schedule, and keep what comes back as a
versioned artifact in a bucket you own.

Two details are worth calling out because they are easy to get wrong by hand.

**Credentials go in on stdin, never in argv.** The request is built as a curl
config file passed on standard input. A token in a command line is visible in
`ps` to every user on the box and frequently ends up in shell history and process
logs. This is the sort of thing that is obvious once stated and almost never done
in a homemade script.

**Every fetch is a separate artifact with its own size and checksum.** So "the
export has been 4 MB every night for a month and last night it was 900 bytes" is
a question something can answer, rather than a thing you find out later.

<Stats>
  <Stat value="1" unit="flag" label="Between an archive and a folder of error pages" />

  <Stat value="0" label="Tokens that appear in argv" />

  <Stat value="1" label="Comparison that catches a login page" />
</Stats>

## When this is the wrong tool [#when-this-is-the-wrong-tool]

Being clear about the boundary: this fetches a URL. It is not a crawler.

If you want a whole site, its assets and its link structure, you want `wget --mirror` or a purpose-built archiver, and you want to think about robots.txt and
about how much of someone else's bandwidth you are entitled to. A web source is
for the specific URL that returns the specific thing you need.

<Aside title="The check worth adding today" tone="accent">
  Whatever fetches a URL for you right now, add `--fail` to it and then look at the
  last thirty files it produced. If they are all within a few percent of each other,
  you are fine. If one of them is suspiciously small, open it. That file is usually
  an error page, and it is usually been there a while.
</Aside>

## The pattern underneath [#the-pattern-underneath]

A URL is a source like any other. What makes it a backup rather than a download
is the same list as everywhere else: it happens on a schedule without anyone
remembering, it keeps versions, it notices when the output stops looking like
itself, and it lands somewhere the thing you are archiving cannot reach.

The curl is one line. The rest is why this is a category.
