Category: Operational Resilience Tags: continuous validation, chaos engineering, SRE, DR testing, RTO, resilience testing, error budgets

The cleanest disaster recovery test I ever signed off on was a lie by August. We ran it in March: a full failover of a tier-one platform to the secondary region, executed against the runbook, restored inside its four-hour objective, evidence captured, auditors satisfied, a green row on the dashboard. Five months later the same platform lost its primary region for real, and the failover did not complete in four hours, or in eight. The postmortem reasons were not exotic. Since March, a new dependency had been added to the recovery path and never tested. A configuration change had quietly left the secondary region short of the capacity it needed at peak. A credential the runbook relied on had rotated. The March test had not been wrong. It had described a system that no longer existed by the time we needed it, which is the same thing as being wrong, just with better paperwork.

This is the structural defect in point-in-time resilience testing: a test certifies the system as it was on the day you ran it, and the system begins drifting away from that state the moment the test ends. Every deploy, every configuration change, every new dependency, every capacity assumption that quietly ages moves your real recovery capability away from the number on the certificate. The annual test does not measure your resilience. It measures your resilience once, and then you spend the next twelve months slowly losing the right to cite the result. The interval between tests is not neutral time. It is where confidence rots, invisibly, until an incident audits it for you.

What SRE Already Solved

The useful news is that another discipline has already worked this out, and you very likely employ people who practice it. Site reliability engineering does not certify a service as reliable once a year. It defines what reliable means as a measurable objective, an SLO, and then measures continuously against an error budget that every degradation spends in real time. Reliability, in that world, is not an event you pass; it is a property you watch. I have written before about making the pivot to an SRE-driven organization, and this is the sharpest place the two worlds meet: resilience and reliability are the same question asked at different blast radii, and SRE answered it with continuous measurement while traditional BC/DR was still booking a conference room for next spring. The resilience program does not need to invent continuous validation from nothing. It needs to borrow the mechanism SRE already runs on and point it at recovery instead of uptime.

Chaos Engineering Is a Control Test, Not a Stunt

The word that scares executives out of this is "chaos." It sounds like breaking production for sport, and the early coverage of the discipline did it no favors. But chaos engineering, done properly, is the most disciplined form of control testing there is, and it is exactly the argument I made in a control that isn't tested is a hope, lifted out of the GRC register and dropped into the running system. A real experiment is hypothesis-driven: we believe that if this instance dies, traffic reroutes and the customer sees nothing. Then you kill the instance, in a controlled window, with a blast radius you have bounded in advance and an abort switch in your hand, and you learn whether the belief was true. A control you assert is a hope; a control you inject a fault against and watch hold is evidence. The failure modes this surfaces (the dependency nobody documented, the retry storm, the failover that needs a human who happens to be asleep) are precisely the ones that turn a four-hour objective into an eight-hour incident. Far better to meet them on a Tuesday afternoon you chose than at 2 a.m. you did not.

Validate the Path, Not the Plan

Continuous validation changes what you test, not just how often you test it. The annual DR test validates a plan, which is to say a document. Continuous validation targets the recovery path: the actual sequence of systems, dependencies, and people that has to work for recovery to happen, with the checking concentrated on the parts most likely to have drifted. The recovery path runs through your vendors, so the fourth-party dependencies I wrote about in the vendor behind your vendor belong inside the validation scope: a failover that assumes a SaaS provider will be there is a hypothesis about someone else's system, and hypotheses get tested. The runbook is itself a validation target, because a plan you cannot execute from under pressure is the knowledge-retrieval failure I described in GRC as a knowledge management discipline in denial: validate that the on-call engineer, not the runbook's author, can find it and follow it. And the thing you are ultimately validating is the number you promised, the tested recovery time that has to sit inside your contractual clock, which is the entire argument of your RTO and your SLA have never met. Continuous validation is how that tested number stays true in the long gaps between the audits that ask for it.

The Metric Moves From Pass to Freshness

Adopt this and your headline metric has to change, because "we passed the annual DR test" stops meaning anything the day after you pass it. The measure that survives contact with reality is coverage-weighted freshness: what fraction of your critical services has a validated recovery recent enough to still be true? A program that validates its tier-one services monthly and its long tail quarterly, through a mix of automated failover checks, injected faults, and targeted exercises, scores worse on the old "annual test completed" line and far better on the real one, which is the same trade I argued for on a different clock in reassessment cadence should be earned. Cadence follows criticality: the services that would hurt most get validated most, and a validation that goes overdue on a tier-one service is not a paperwork slip, it is a live gap in the only assurance that matters. The annual full-scale test does not disappear in this model. It becomes the anchor and the audit artifact, the deep all-hands rehearsal you still run once a year, sitting on top of a continuous stream of smaller checks that keep the green from going stale between rehearsals.

Go back to that March test. In a continuous-validation program, the new dependency on the recovery path would have failed a check the week it landed, not five months later in a live outage. The capacity drift would have shown up in a routine failover probe. The rotated credential would have broken a scheduled test loudly, on a day when nobody was panicking. None of this requires heroics or a new budget line the size of a product team; it requires treating recovery the way SRE already treats reliability, as a property you measure continuously rather than an event you pass annually. The certificate on the wall tells you the system recovered once. Your customers, and your regulators, are asking whether it will recover today, and only continuous validation can answer in the present tense.

PivotRisk is a practitioner-led governance, risk and resilience practice. Its work comes from 20+ years building GRC, continuity and security programs across global fintech, SaaS and enterprise technology, including continuous resilience validation built on SRE, resilience-engineering and chaos-engineering principles.

Testing as a programme, not an event

DORA codifies exactly this shift: a digital operational resilience testing programme, run continuously and risk-based, not an annual box-tick. The DORA Readiness Assessment scores your testing-programme maturity alongside the rest of the regulation and shows you where point-in-time testing leaves a gap an examiner will find.

Get the DORA Readiness Assessment