Plenty of things worth keeping are only available over HTTP. An export endpoint on a SaaS you do not control. A published dataset that changes. A rendered report. A partner's feed that they will regenerate and overwrite next Tuesday.
Fetching one is trivial. Everything after that is the work.
Start with the curl that does not lie to you
curl --fail --location --silent --show-error \
--retry 3 --retry-delay 5 \
-H "Authorization: Bearer $TOKEN" \
-o export.json \
https://api.example.com/v1/export--fail is the one people leave out, and it is the one that matters. Without it
curl exits zero on a 404 or a 500 and cheerfully writes the error body to your
file. Your archive fills up with plausibly-sized HTML pages that say "Sign in".
--location follows redirects, which is how an endpoint that moved keeps working.
--retry covers the transient 502 that would otherwise lose you a night.
The four things that go wrong afterwards
The response quietly changes shape. The endpoint starts returning a login page, or an empty array, or a truncated body. All three produce a file. None of them produce an error. The only defence is comparing this run's size against the last one, which means something has to be recording the sizes.
The token expires. Usually on a weekend, usually silently, and usually discovered when someone needs the data.
Nobody keeps the versions. curl -o export.json overwrites. If you want last
Tuesday's copy, you needed to have decided that last Tuesday.
It lives on someone's laptop. The most common home for a script like this is a machine that gets replaced.
What a web source does with it
We run essentially the curl above, on a schedule, and keep what comes back as a versioned artifact in a bucket you own.
Two details are worth calling out because they are easy to get wrong by hand.
Credentials go in on stdin, never in argv. The request is built as a curl
config file passed on standard input. A token in a command line is visible in
ps to every user on the box and frequently ends up in shell history and process
logs. This is the sort of thing that is obvious once stated and almost never done
in a homemade script.
Every fetch is a separate artifact with its own size and checksum. So "the export has been 4 MB every night for a month and last night it was 900 bytes" is a question something can answer, rather than a thing you find out later.
1flag
Between an archive and a folder of error pages
0
Tokens that appear in argv
1
Comparison that catches a login page
When this is the wrong tool
Being clear about the boundary: this fetches a URL. It is not a crawler.
If you want a whole site, its assets and its link structure, you want wget --mirror or a purpose-built archiver, and you want to think about robots.txt and
about how much of someone else's bandwidth you are entitled to. A web source is
for the specific URL that returns the specific thing you need.
The pattern underneath
A URL is a source like any other. What makes it a backup rather than a download is the same list as everywhere else: it happens on a schedule without anyone remembering, it keeps versions, it notices when the output stops looking like itself, and it lands somewhere the thing you are archiving cannot reach.
The curl is one line. The rest is why this is a category.