Web
Fetch what a URL serves, on a schedule.
View as MarkdownFetches a URL and stores the response body. Available as both a local and a cloud backup.
It is the connector for anything that already has an export endpoint: a SaaS API that returns a dump, an internal reporting service, a status feed you want a record of.
What is taken
The response body, exactly as served. Nothing is parsed, rewritten or followed.
| Included | Not included |
|---|---|
| The body of the final response | Linked pages, assets or scripts |
| Response headers, beyond naming the file | |
| Anything requiring a browser to render |
This is a fetch, not a crawler. One URL, one request, one artifact. If you need many pages,
that is a script source calling the crawler of your choice.
Fields
| Key | Required | Notes |
|---|---|---|
url | Yes | Must start with https:// |
method | No | Defaults to GET |
headers | No | One Name: value per line |
An http:// URL is refused, not silently upgraded. Someone who typed http:// meant it,
and rewriting the scheme would hide that the fetch was in the clear. Fix the URL, or if the
endpoint really is plaintext-only, fetch it through a
script source where the choice is visible in your own code.
On a cloud backup, url and method are public and headers are secret, because that is
where the bearer token lives.
backups:
"<backup-uuid>":
source_type: web
source:
url: https://api.internal/export
headers: |
Authorization: Bearer <token>
Accept: application/jsonbackups:
- name: vendor-export
kind: cloud
source_type: web
schedule: "0 4 * * *"
source:
url: https://api.vendor.com/v1/export
headers: |
Authorization: Bearer <token>A value must not contain a line break, and one is refused rather than accepted. A newline inside a header value would let the value become a second header, which is a request smuggling shape we would rather not have.
Redirects and schemes
Redirects are followed, but only to http and https. A redirect to file:// cannot
turn a web backup into a filesystem read, on either kind. The starting URL must be https://;
a redirect is allowed to land on http://, because that is the server's choice and not
something the config claimed.
The artifact is named from the response's Content-Disposition header, falling back to the
last path segment of the final URL after redirects.
Failures
An HTTP error status fails the run. A 500 from the source is not something to archive as though it were data.
| Situation | Result |
|---|---|
2xx | The body is stored |
4xx, 5xx | The run fails, with the status in the error |
| Redirect to a non-HTTP scheme | The run fails |
| Connection refused, DNS failure, TLS error | The run fails and is retried |
A 200 with an error message in the body is stored as a successful backup. Some APIs answer
that way. If yours does, use a script source that checks
the payload and exits non-zero, rather than discovering it during a restore.
Authentication
Whatever the endpoint accepts as a header: a bearer token, an API key, basic auth.
The header lives with the source credentials, which means it lives on your worker for a local backup and in our vault for a cloud one. Prefer a read-only, export-scoped token over a general-purpose one, and rotate it by reconfiguring the backup.
A token that expires is the most common cause of a web backup that worked for months and then
stopped. The failure is a 401 in the run's error text.
Sizing and timeouts
The whole body is written before the run moves on, so the artifact is exactly the size of the response.
A slow endpoint is fine, up to a point: the fetch runs inside the run's step timeout, and
progress is reported by byte count so a large download is visibly alive rather than hung. An
endpoint that streams for hours is better served by a script source you control the
timeouts on.
Restoring
There is nothing to restore into. The artifact is the file the endpoint served, and what you do with it is up to you.
sctl restore file <artifact-id> --path ./export.json