---
title: "RPO and RTO for people who just have a cron job"
description: "Two acronyms that sound like enterprise procurement and are actually just two numbers you already have. Here is how to work out yours in an afternoon, and why the second one is usually a guess."
url: "https://saved.sh/blog/rpo-and-rto-for-a-cron-job"
date: "2026-08-08"
author: "saved.sh"
tag: "Engineering"
---

RPO and RTO show up in vendor decks and compliance questionnaires, wrapped in
enough language to make them sound like something you buy rather than something
you measure.

They are two numbers. You already have both. You probably have not written them
down, and the second one is almost certainly wrong.

## The two numbers [#the-two-numbers]

**RPO, recovery point objective: how much data you are willing to lose.** It is a
window of time, measured backwards from the incident. Your RPO is set entirely by
how often your backups run.

**RTO, recovery time objective: how long you are willing to be down.** Measured
forward from the moment someone decides to restore. Your RTO is set by how long
the restore actually takes, plus everything that happens before anyone starts it.

<Figure caption="RPO looks backwards from the incident and is set by your schedule. RTO looks forwards and is set by your restore.">
  <RunHeatmap />
</Figure>

## Working out your real RPO in one minute [#working-out-your-real-rpo-in-one-minute]

If your cron runs at 03:00 daily, your RPO is 24 hours. An incident at 02:55
loses almost a full day of writes.

That is the arithmetic everyone does. It is also optimistic, because it assumes
last night's run succeeded. Your **effective RPO is the age of your most recent
restorable backup**, which is a different number:

| Situation                                         | Stated RPO | Effective RPO                    |
| ------------------------------------------------- | ---------- | -------------------------------- |
| Nightly job, all green                            | 24 hours   | 24 hours                         |
| Nightly job, failing quietly for 6 days           | 24 hours   | 7 days                           |
| Nightly job, succeeding but truncated for 3 weeks | 24 hours   | Unbounded, you have no good copy |

The gap between the two columns is the whole reason run records exist. Without a
record of what each run produced, you know your schedule, which is a statement of
intent, and not your RPO, which is a fact.

<Stats>
  <Stat value="24" unit="hours" label="What most teams believe their RPO is" />

  <Stat value="1" label="Failed run needed to make that false" />

  <Stat value="0" label="Alerts a silent cron failure produces" />
</Stats>

## Your RTO is a guess until you have timed it [#your-rto-is-a-guess-until-you-have-timed-it]

Nearly every RTO figure quoted internally is the time the restore command takes,
measured once, on a good day, by the person who wrote it.

The real clock starts earlier and runs longer:

1. **Detection.** Someone notices. On a bad day this is hours.
2. **Decision.** Someone decides restoring is correct, which usually needs a
   second person who is asleep.
3. **Locating.** Which artifact, from when, in which bucket, under which key.
4. **Retrieval.** Download time, which people forget scales with size and their
   own bandwidth. 200 GB is not instant.
5. **Restore.** The number everybody quotes.
6. **Verification.** Confirming it worked before you point traffic at it.
7. **Reconciliation.** Replaying whatever happened during steps 1 through 6.

Steps 1, 2 and 7 are usually larger than step 5, and none of them appear in a
vendor's RTO claim, because none of them are the vendor's.

## How to actually measure both [#how-to-actually-measure-both]

An afternoon, once:

**For RPO**, look at your most recent *verified* backup, not your most recent
backup attempt. If you cannot tell the difference from your tooling, that is the
finding. Age of that artifact is your effective RPO.

**For RTO**, run the restore. Start a timer when you decide to, not when you type
the command. Restore into a scratch environment, verify row counts, and stop the
timer when you would have been comfortable sending traffic to it. Whatever number
comes out is your RTO, and it will be larger than what is written down.

<Aside title="Do the cheap half first" tone="accent">
  If you only do one of these, do the RTO test. RPO can be improved by changing a
  cron schedule, which takes a minute. RTO is improved by discovering the six
  things that break during a restore, and you can only find those by doing one.
</Aside>

## What we do with these two numbers [#what-we-do-with-these-two-numbers]

We do not set your RPO. Your schedule does, and that is your call.

What we do is make the effective number match the stated one, which is where
these usually diverge. Every run leaves a record: what ran, what it produced, how
large it was against previous runs, where each copy went, and what failed. A
missing or anomalous run is visible rather than silent, so the age of your last
good artifact is a thing you can read rather than a thing you assume.

For RTO, our contribution is smaller and more specific. The restore path is
documented, uses standard tools, and does not require us to be reachable, which
removes one of the steps above from your critical path entirely.

## The summary [#the-summary]

Your RPO is your schedule, corrected for the runs that failed. Your RTO is a
guess until you time a real restore, and it is longer than you think, mostly
because of the parts that have nothing to do with technology.

Write both down. They are the two numbers that determine what a bad day costs
you, and neither of them needs a procurement process to find out.
