Category: Operating Models Author: Cody Swidler Tags: knowledge management, runbooks, document control, policy administration, audit readiness, BC/DR, evidence management

During a datacenter failover a few years back, the recovery team did everything right except find the plan. There were four copies of the DR runbook: one in SharePoint — unreachable, because SharePoint was in the affected environment, a circular dependency nobody had noticed until that moment — one in Confluence, two major versions stale; one on the recovery lead's laptop, annotated and current, but he was on a plane; and one printed in a binder in an office nobody had visited since the lease downsized. The team recovered the environment on memory and group chat, in roughly double the tested time, and the retrospective dutifully recorded the miss as a "documentation issue." It wasn't a documentation issue. The documentation existed — four times. It was a knowledge management failure: the organization knew how to recover, and could not retrieve what it knew at the only moment the knowing mattered.

I've come to believe this is the quiet defect underneath most GRC and resilience programs. Strip the acronyms off and look at the actual assets: policies, standards, procedures, response plans, runbooks, risk registers, control descriptions, audit evidence, contact trees, vendor files. Every one of them is knowledge — captured, versioned, owned, decaying knowledge. And yet almost no GRC program has anyone responsible for knowledge management as a discipline. We have writers everywhere and librarians nowhere. The result is a program that produces documents at industrial scale and retrieves the right one at the worst possible moment approximately by luck.

A Runbook Is Knowledge With a Half-Life

The reframe matters most in BC/DR, because that's where retrieval happens under the least forgiving conditions. A response plan is not a compliance artifact that happens to contain instructions; it is operational knowledge with a half-life, and the half-life is short. Contact names go stale the day someone resigns. Architecture diagrams expire with every migration. The failover sequence documented in January describes a topology that stopped existing in March. A plan is only as real as its most recently verified fact, which is why exercises are best understood as knowledge validation events — the tabletop isn't just testing the people, it's testing whether the written knowledge still matches the world. And retrieval is part of the asset, not an afterthought: knowledge you cannot reach during the incident it describes is indistinguishable from knowledge you don't have. That's the argument for the unfashionable battle box — the offline, out-of-band copy of the plans, contacts, and credentials-access procedures that survives the loss of the systems they describe. It looks like paranoia until the wiki is inside the blast radius, at which point it looks like the only KM decision anyone got right.

Policy Administration Is Document Control Wearing a Suit

The same reframe explains why policy administration — the unglamorous machinery of owners, approvers, review dates, version histories, and distribution — is perpetually underbuilt. Organizations treat policy as a writing problem, so they invest in authorship and neglect custody. But nearly everything that goes wrong with policy is a custody failure. The orphaned policy whose owner left two reorgs ago. The "current" version that differs depending on which drive you searched. The commitment nobody can trace to an author, which is exactly how the Good Idea Fairy's sentences achieve tenure — unowned knowledge can't be challenged, because there's nobody to ask what it meant. A real policy administration function is a knowledge lifecycle: every governing document has one system of record, one accountable owner, a review date that triggers actual review, a version history that explains what changed and why, and a retirement path so superseded knowledge leaves circulation instead of lingering as a trap. None of this is intellectually hard. It's just work that belongs to a discipline most programs never staffed.

An Audit Is a Knowledge Retrieval Exam

Here's the part that should get budget attention: every certification and audit your organization undergoes is, mechanically, a retrieval test. ISO 27001 devotes an entire clause to documented information — current versions, controlled changes, availability where needed — and an experienced auditor opens with the document register because it reveals in ten minutes whether the management system is real. SOC 2 fieldwork is a sustained exercise in producing the right evidence, for the right period, in the right version, on request. The dreaded audit-prep scramble — three weeks of hunting for screenshots, chasing document owners, and discovering the access review evidence lives in a departed contractor's email — is not an audit problem. It is the annual invoice for the KM function you didn't build, paid in overtime. Programs with genuine knowledge management don't do audit prep in any meaningful sense; evidence lands in a known location with metadata when the control operates, documents carry their approval history with them, and the auditor's request list becomes a series of lookups instead of investigations. "Audit-ready" is not a state of virtue. It's a retrieval latency.

The Library, Not the Pile

What does good look like? Less than you'd fear. One system of record per document class — not one tool for everything, but one authoritative home for each thing, with every other copy explicitly a copy. Metadata as a requirement, not a courtesy: owner, approver, review date, version, audience, and the process or system the knowledge describes, so documents can be found by what they're for rather than by what someone remembered to title them. Review cycles driven by the metadata, with staleness surfaced as a metric — percent of response plans validated in the last exercise cycle, percent of policies inside their review window, count of orphaned documents — reported alongside your other program KPIs, because a stale runbook is a risk exposure wearing a filename. Retrieval paths tested under the conditions that matter, including the ugly ones where the network is gone. And structure over prose wherever the knowledge will be consumed by a machine or a panicked human: a contact tree as a maintained dataset rather than a table pasted into a PDF, recovery steps as discrete, testable units rather than paragraphs. This is the same trajectory I described in making risk data computable — every AI ambition your GRC roadmap contains quietly assumes a knowledge base that is current, structured, and authoritative, which is to say it assumes the KM discipline nobody funded. The organizations that get agents summarizing their runbooks will be the ones whose runbooks were worth summarizing.

The failover story ended the way these stories should: the program hired no one and bought nothing, but it appointed custodians, collapsed four copies into one system of record plus one battle box, put review dates and owners on every plan, and started grading exercises partly on retrieval — could the on-call engineer, not the author, find and follow the current version? Two years later an auditor described their documentation as the cleanest she'd seen that year, and the resilience lead told me the compliment felt strange, because none of it had been done for her. That's the tell that the discipline is working: the audit gets easy as a side effect. The knowledge was managed for the 2 a.m. incident — the auditor just happened to benefit from arriving at 2 p.m.

Cody Swidler is the founder of PivotRisk and Head of Platform Resiliency at Apex Fintech Solutions. He has built and scaled GRC, resilience, and risk programs across Microsoft, Twilio, Box, Zayo, and Miro.

Cody Swidler is the founder of PivotRisk and Head of Platform Resiliency at Apex Fintech Solutions. He has built and scaled GRC, resilience, and risk programs across Microsoft, Twilio, Box, Zayo, and Miro.

Plans built to be found and followed

The Business Continuity Plan Template structures the plan as maintainable knowledge — owned sections, contact trees, and a battle box checklist for the copy that survives losing your systems — and the Tabletop Exercise Playbook turns each exercise into a test of whether the written plan still matches reality.

Browse the Templates