Category: Operational Resilience Author: Cody Swidler Tags: RTO, SLA, recovery objectives, contractual commitments, BIA, impact tolerance, customer trust
The strangest incident retrospective I've ever attended was for a recovery that worked. A core platform went down hard, the team executed the runbook, and the service was restored inside its 24-hour recovery time objective — a genuine operational success, and the room knew it. Then someone from legal joined late, holding a printout of the master service agreement for the company's largest customer. It committed to restoration of service within four hours, with escalating credits after that, and there was similar language — negotiated separately, over several years, by several different account teams — in the contracts of eleven other enterprise customers. The company had just completed a successful recovery and a dozen contractual breaches simultaneously, with the same outage, on the same clock. The RTO had been met. The SLA had never been consulted. The two numbers described the same system, and they had never been in the same room.
That's the reconciliation problem in one scene. The recovery time objective is an internal promise: born in a BIA workshop, owned by the resilience team, reviewed annually, and tested — if you're disciplined — against what the infrastructure can actually do. The service level agreement is an external promise: born in a deal negotiation, owned by sales and legal, priced under commercial pressure, and tested by nothing except an incident. They are written in different documents, by different functions, in different units, on different calendars. At most organizations, nobody's job description includes making them agree.
Two Promises, Two Authors, and a Deal Desk
Understand how each number gets made and the divergence stops being surprising. The RTO comes from impact analysis: how long can the business tolerate this process being down before the damage becomes unacceptable? It's an engineering-facing number, and its natural failure mode is aspiration — a tier label picked in a workshop and never validated against a tested recovery. The SLA comes from a negotiation in which recovery capability is not in the room but revenue is. An enterprise prospect's procurement team asks for four-hour restoration because their last vendor agreed to it; the account executive, staring at a quarter-end close date, agrees; legal wordsmiths the credits clause; and nobody routes the term to anyone who knows the tested recovery time is measured in days. The commitment isn't reckless because anyone was reckless. It's reckless because the organization has no path by which the person signing the promise would ever learn what the promise costs. Multiply that by years of deals, each negotiated independently, and your real recovery obligation is not your RTO — it's the tightest SLA any account team has ever signed, which is a number most resilience programs have never seen.
The Two Numbers Don't Even Share Units
Reconciliation is harder than comparing two durations, because the promises are usually denominated differently. SLAs love availability percentages — 99.9%, "three nines" — and percentages are annual accounting, not incident accounting. A 99.9% commitment tolerates about eight and three-quarter hours of downtime a year, which sounds compatible with an eight-hour RTO until you notice it says nothing about how that downtime arrives. One clean eight-hour outage can be simultaneously SLA-compliant for the year and a catastrophic breach of every per-incident restoration clause buried in the same contract. Meanwhile the RTO speaks only in per-incident terms and says nothing about frequency: an environment that fails monthly but recovers in an hour beats its RTO every time while burning through the availability budget by June. And the same mismatch runs through the data dimension — your RPO is an internal tolerance for data loss, while the contract may promise "no loss of customer data," an absolute, written by someone who has never heard of a replication lag. Until you translate every commitment into the same units — per-incident restoration time, per-incident data loss, and annual availability, side by side — you don't actually know what you've promised, only what each document sounds like.
The Stack Beneath You Made Its Own Promises
The third reconciliation isn't between your numbers — it's between your numbers and your suppliers'. Your four-hour RTO is not a property of your intentions; it's a property of the slowest critical dependency in the recovery path. If the platform you'd restore runs on a SaaS vendor whose own SLA is 99.5% with next-business-day support, then your four-hour commitment quietly assumes a counterparty who has promised you nothing of the kind. SLA math does not compose upward: stack three dependencies at 99.9% each and your theoretical ceiling is already below any of them individually — and their restoration clauses don't inherit your urgency. This is the same blindness I wrote about in fourth-party risk, expressed in hours instead of org charts, and it's precisely the chain regulators are now pulling: DORA's testing and register requirements exist because supervisors stopped believing that firms knew whether their contractual promises survived contact with their vendors' contractual promises. For any critical service, there is a vendor floor — the recovery time your dependencies will actually commit to — and no RTO below that floor is real, no matter what the BIA workshop voted.
Reconciliation Is a Table, Not a Treaty
The fix does not require a cross-functional transformation program; it requires one table with a row per critical service and an owner who maintains it. Five columns. The tightest customer commitment across every signed contract — which means someone must actually mine the contracts, because the number lives in deviations from the standard terms, not in the template. The regulatory or impact tolerance, where one applies. The stated RTO. The last tested recovery time — the achieved number, not the objective, in the spirit of a control that isn't tested is a hope. And the vendor floor beneath the service. The whole discipline is one inequality, read left to right: tested recovery ≤ RTO ≤ tightest SLA, with the vendor floor underneath the tested number. Every service where the inequality holds is a promise you're entitled to keep making. Every service where it breaks is a decision, and there are only three: invest until the tested number clears the promise, renegotiate the promise down to what you can prove, or knowingly carry the gap as an accepted risk with an executive's name on it — priced, not discovered by legal in a war room.
And then close the loop where the divergence starts: the deal desk. Non-standard availability or restoration language in any new contract routes to the resilience function for sign-off before signature, the same way non-standard payment terms route to finance. This is not bureaucracy; it's the customer trust function working in its highest-leverage position — before the promise exists. A sales team that can say "we can commit to eight hours, contractually, because we test to five" closes deals on evidence. That's a stronger pitch than four hours of fiction, and dramatically cheaper than the credits.
Back to that retrospective. The fix wasn't heroic engineering — the recovery was already good. It was the table: twelve contracts re-papered over eighteen months to restoration language the tested capability actually cleared, one genuinely strategic customer whose four-hour term was kept and funded with a dedicated warm standby, and a deal-desk gate so the thirteenth contract never got signed blind. The RTO and the SLA finally met. Like most estranged relatives, they turned out to have less in common than everyone assumed — which is exactly why someone has to formally introduce them, in writing, before an incident does it for you.
Cody Swidler is the founder of PivotRisk and Head of Platform Resiliency at Apex Fintech Solutions. He has built and scaled GRC, resilience, and risk programs across Microsoft, Twilio, Box, Zayo, and Miro.
Put the promises in one table
The Business Impact Analysis template scores your processes, derives the recovery objectives, and exposes the gap between what you've promised and what you can prove — the reconciliation table this article describes, built and formula-driven.
Get the BIA Template