Category: Operational Resilience Author: Cody Swidler Tags: RTO, SLA, recovery objectives, contractual commitments, force majeure, SRE, error budget, BIA, customer trust
The strangest incident retrospective I've ever attended was for a recovery that worked. A core platform went down hard, the team executed the runbook, and the service was restored inside its 24-hour recovery time objective — a genuine operational success, and the room knew it. Then someone from legal joined late, holding a printout of the master service agreement for the company's largest customer. It committed to restoration of service within four hours, with escalating credits after that, and there was similar language — negotiated separately, over several years, by several different account teams — in the contracts of eleven other enterprise customers. The company had just completed a successful recovery and a dozen contractual breaches simultaneously, with the same outage, on the same clock. The RTO had been met. The SLA had never been consulted. The two numbers described the same system, and they had never been in the same room.
That's the reconciliation problem in one scene. The recovery time objective is an internal promise: born in a BIA workshop, owned by the resilience team, reviewed annually, and tested — if you're disciplined — against what the infrastructure can actually do. The service level agreement is an external promise: born in a deal negotiation, owned by sales and legal, priced under commercial pressure, and tested by nothing except an incident. They are written in different documents, by different functions, in different units, on different calendars. At most organizations, nobody's job description includes making them agree.
Two Promises, Two Authors, and a Deal Desk
Understand how each number gets made and the divergence stops being surprising. The RTO comes from impact analysis: how long can the business tolerate this process being down before the damage becomes unacceptable? It's an engineering-facing number, and its natural failure mode is aspiration — a tier label picked in a workshop and never validated against a tested recovery. The SLA comes from a negotiation in which recovery capability is not in the room but revenue is. An enterprise prospect's procurement team asks for four-hour restoration because their last vendor agreed to it; the account executive, staring at a quarter-end close date, agrees; legal wordsmiths the credits clause; and nobody routes the term to anyone who knows the tested recovery time is measured in days. The commitment isn't reckless because anyone was reckless. It's reckless because the organization has no path by which the person signing the promise would ever learn what the promise costs. Multiply that by years of deals, each negotiated independently, and your real recovery obligation is not your RTO — it's the tightest SLA any account team has ever signed, which is a number most resilience programs have never seen.
And there is one more asymmetry, the one that quietly does the most damage: only one of these promises is binding. The SLA is in the contract — signed, countersigned, credits attached, enforceable. The RTO is in none of them. It lives in due-diligence questionnaires and BIA workshops, gets discussed at length during every enterprise sales cycle, and then binds precisely no one. This produces a subtle and expensive substitution: teams pour their attention into the number they own and control — the RTO — and forget it was never the number that governs the penalty. And if your RTO is set equal to or looser than your SLA, then meeting your RTO in a serious incident still breaches your SLA, because detection, declaration, and failover all ran the clock while you were busy hitting the target that doesn't count. Optimizing the RTO in isolation doesn't protect the contract — it guarantees you miss it in exactly the incidents that matter most.
The SLA Has a Trapdoor You Only Find by Falling Through It
Before you conclude the SLA is simply the hard number to plan around, read the rest of the clause. Almost every enterprise SLA is caveated — force majeure language, disaster carve-outs, "conditions beyond reasonable control" — that suspends the credits when an outage stems from a qualifying event. On paper this looks like relief: surely a true disaster won't count against us. But notice when that question actually gets answered. Not at signing. Not during the incident. After — when both sides' lawyers sit down with the same paragraph and reach opposite conclusions about whether your outage was a force-majeure disaster or just a bad Tuesday. In my experience these debates resolve only in hindsight, and the outcome turns on facts you won't know until you're in them: the root cause, whether you followed your own plan, whether a prudent operator could have prevented it. You cannot build a resilience strategy on a clause whose meaning is decided after the fact by opposing counsel. The only safe assumption is that the trapdoor is welded shut — plan to honor the SLA as written, and treat any force-majeure relief you eventually win as an unexpected refund, never as a line in the plan.
The Two Numbers Don't Even Share Units
Reconciliation is harder than comparing two durations, because the promises are usually denominated differently. SLAs love availability percentages — 99.9%, "three nines" — and percentages are annual accounting, not incident accounting. A 99.9% commitment tolerates about eight and three-quarter hours of downtime a year, which sounds compatible with an eight-hour RTO until you notice it says nothing about how that downtime arrives. One clean eight-hour outage can be simultaneously SLA-compliant for the year and a catastrophic breach of every per-incident restoration clause buried in the same contract. Meanwhile the RTO speaks only in per-incident terms and says nothing about frequency: an environment that fails monthly but recovers in an hour beats its RTO every time while burning through the availability budget by June. And the same mismatch runs through the data dimension — your RPO is an internal tolerance for data loss, while the contract may promise "no loss of customer data," an absolute, written by someone who has never heard of a replication lag. Until you translate every commitment into the same units — per-incident restoration time, per-incident data loss, and annual availability, side by side — you don't actually know what you've promised, only what each document sounds like.
Set the RTO Inside the SLA — On Purpose
Here is the reconciliation that actually resolves the conflict, and the industry already worked it out: the recovery objective has to be strictly smaller than the commitment it protects. This is the discipline underneath Google's SRE practice and the one I keep pushing resilience teams toward in the SRE pivot for resilience — you don't set the RTO to what you hope to achieve and then pray it lands under the SLA. You set it deliberately below the SLA, with margin, so that a full disaster-recovery execution — detection, declaration, failover, validation, the messy human first minutes and all — still finishes before the contractual clock runs out. The SLA is the ceiling; the RTO is the ceiling minus the margin you need to survive a bad day. When the RTO is engineered inside the SLA, meeting your recovery objective and honoring your customer commitment become the same event, even in a disaster — which is the whole point of a recovery target: clearing it should be what keeps you compliant. When the RTO sits at or above the SLA, the two diverge at the worst possible moment, and you get the retrospective this article opened with — a clean recovery and a dozen breaches, on the same clock.
The Stack Beneath You Made Its Own Promises
Setting the RTO inside the SLA is only possible if the stack beneath you allows it — so the last reconciliation isn't between your numbers, it's between your numbers and your suppliers'. Your four-hour RTO is not a property of your intentions; it's a property of the slowest critical dependency in the recovery path. If the platform you'd restore runs on a SaaS vendor whose own SLA is 99.5% with next-business-day support, then your four-hour commitment quietly assumes a counterparty who has promised you nothing of the kind. SLA math does not compose upward: stack three dependencies at 99.9% each and your theoretical ceiling is already below any of them individually — and their restoration clauses don't inherit your urgency. This is the same blindness I wrote about in fourth-party risk, expressed in hours instead of org charts, and it's precisely the chain regulators are now pulling: DORA's testing and register requirements exist because supervisors stopped believing that firms knew whether their contractual promises survived contact with their vendors' contractual promises. For any critical service, there is a vendor floor — the recovery time your dependencies will actually commit to — and no RTO below that floor is real, no matter what the BIA workshop voted. And if that vendor floor sits above your tightest SLA, no amount of internal engineering closes the gap — you are choosing between a different vendor, a renegotiated customer commitment, or a breach you have already scheduled.
Reconciliation Is a Table, Not a Treaty
The fix does not require a cross-functional transformation program; it requires one table with a row per critical service and an owner who maintains it. Five columns. The tightest customer commitment across every signed contract — which means someone must actually mine the contracts, because the number lives in deviations from the standard terms, not in the template. The regulatory or impact tolerance, where one applies. The stated RTO. The last tested recovery time — the achieved number, not the objective, in the spirit of a control that isn't tested is a hope. And the vendor floor beneath the service. The whole discipline is one inequality, read left to right: tested recovery ≤ RTO < tightest SLA, with the vendor floor underneath the tested number. That middle sign is strict on purpose — the RTO has to sit inside the SLA, not merely touch it, or the margin a real incident consumes doesn't exist. Every service where the inequality holds is a promise you're entitled to keep making. Every service where it breaks is a decision, and there are only three: invest until the tested number clears the promise, renegotiate the promise down to what you can prove, or knowingly carry the gap as an accepted risk with an executive's name on it — priced, not discovered by legal in a war room.
And then close the loop where the divergence starts: the deal desk. Non-standard availability or restoration language in any new contract routes to the resilience function for sign-off before signature, the same way non-standard payment terms route to finance. This is not bureaucracy; it's the customer trust function working in its highest-leverage position — before the promise exists. A sales team that can say "we can commit to eight hours, contractually, because we test to five" closes deals on evidence. That's a stronger pitch than four hours of fiction, and dramatically cheaper than the credits.
Back to that retrospective. The fix wasn't heroic engineering — the recovery was already good. It was the table: twelve contracts re-papered over eighteen months to restoration language the tested capability actually cleared, one genuinely strategic customer whose four-hour term was kept and funded with a dedicated warm standby, and a deal-desk gate so the thirteenth contract never got signed blind. The RTO and the SLA finally met. Like most estranged relatives, they turned out to have less in common than everyone assumed — which is exactly why someone has to formally introduce them, in writing, before an incident does it for you.
Cody Swidler is the founder of PivotRisk and Head of Platform Resiliency at Apex Fintech Solutions. He has built and scaled GRC, resilience, and risk programs across Microsoft, Twilio, Box, Zayo, and Miro.
Put the promises in one table
The Business Impact Analysis template scores your processes, derives the recovery objectives, and exposes the gap between what you've promised and what you can prove — the reconciliation table this article describes, built and formula-driven.
Get the BIA Template