Jump to content

Run a Documentation Audit in a Single Afternoon

Engineering Culture Silas Whitlock

The Danger of Undead Runbooks

An HTTP 200 response masking a retired deployment command represents one of the most dangerous artifacts in software engineering. A working page often hides a broken process. Engineers executing outdated runbooks during an outage waste critical minutes debugging the documentation itself. The solution requires inventorying the docs, checking for detectable failures, prioritizing by operational impact, and assigning an owner to every page worth keeping.

Restrict the initial audit scope to onboarding, deployment, incident, and access documentation, and the inventory phase fits into a single afternoon. Attempting to catalog every architectural decision record and meeting note guarantees the project will stall. Focus entirely on the operational guides that dictate how systems are built, deployed, and recovered.

Audit Scope Priority

Limit the initial pass to incident and deployment guides to guarantee completion within a 3-to-4 hour afternoon window.

Because the initial pass is boxed into one afternoon, tight constraints force ruthless prioritization. You will ignore the sprawling product wikis and focus exclusively on the text that keeps the infrastructure running.

Merging Repositories and Wikis

Engineering knowledge rarely lives in a single system. You must collect Markdown paths from the relevant repositories and page URLs from each wiki space or export available to the team. Permission-limited indexes frequently hide pages from automated discovery tools, requiring manual exports from wiki administrators to build a complete picture.

Build a 14-column tracking sheet structure including URL, system, doc type, and staleness. Add columns for last checked, last meaningful review, owner, fallback owner, impact, link health, ownership, action, due date, and evidence. This structure forces a full evaluation of each artifact.

Deduplication of 301 redirects and cloned wiki spaces prevents redundant work. Flag pages lacking an obvious parent index or link from a maintained entry point as possible orphans. Structure the tracking sheet to explicitly separate unlinked pages from pages that are merely inaccessible to the auditor, preventing false-positive orphan flags. An inaccessible page requires a permissions fix, while a true orphan requires deletion.

Scripting the Discovery Phase

Initially, parsing Markdown ASTs to validate code block syntax seemed viable for detecting broken commands. We abandoned this approach because it flagged too many intentional pseudocode examples. Shift to a simpler pipeline using standard version control utilities.

Use git ls-files -z '*.md' to enumerate tracked Markdown files. Pipe that output into git log -1 --format=%cs -- "$file" to record the latest commit date for each path. Reviewing the Git log documentation provides additional formatting options for extracting author metadata. Edit history serves as a review cue while offering zero proof of technical accuracy. A file touched yesterday might contain a typo that breaks a production deployment.

Image showing script workflow

Build a link checker that extracts destination URLs and resolves relative paths against the source file directory. Check local targets directly. Probe external URLs conservatively, rate-limiting to 2 requests per second to avoid triggering WAF blocks. Treat authentication failures and URL fragments as manual-review cases. Compare the wiki export or API page list with maintained index pages to produce an orphan-candidate queue. Automated discovery remains incomplete when access permissions or dynamically generated navigation obscure pages.

Scoring Operational Blast Radius

Evaluate the collected inventory using a 0 to 8 point total rubric scale based on impact, staleness, link health, and ownership. Score each category from 0 to 2, then sum them. Define each end of the scale in plain language so two reviewers can apply it consistently across different repositories.

Make the impact metric concrete. Score a production rollback guide with a 3-year-old timestamp against a 6-month-old team introduction page and the priority inversion shows up immediately. The rollback guide outranks the intro page entirely due to the damage it could cause during an incident. This scoring rubric breaks down for customer-facing API references where formatting and brand voice carry weight; it is strictly designed for internal operational runbooks.

Force Triage Decisions

Use scores to force decisions instead of building a leaderboard. High-impact pages with low link health require immediate verification.

Fix a dangerous instruction immediately. Verify an uncertain page with a maintainer. Redirect or archive a duplicate. Leave a healthy page alone.

Assigning Genuine Maintenance Authority

Assign ownership to the team responsible for the underlying system. Map ownership to a specific internal routing tag or on-call rotation rather than an individual LDAP alias. Individuals change teams or leave the company, leaving documentation orphaned. A routing tag ensures the current on-call engineer receives the maintenance request.

Define ownership explicitly as the authority to retire a page or accept corrections. Move away from the default assumption that the last Git committer is the current maintainer. The engineer who fixed a typo six months ago does not own the system architecture.

Where no responsible team can be found, flag the page for an explicit keep-or-archive decision. Give pages with orphaned owner fields a 14-day keep-or-archive window to force a resolution. An orphaned owner field is itself an audit finding that requires remediation.

Integrating Checks Into Deployments

Turn the tracking sheet into a small maintenance queue with an owner, next action, and due date. Integrate the documentation check directly into the deployment pipeline rather than relying on calendar reminders. This ensures high-impact pages are re-evaluated when the underlying system changes.

Adding a mandatory documentation check step to the standard pull request template for infrastructure changes prevents drift at the source. When an engineer modifies a Terraform module, the PR template forces them to link the updated runbook. Schedule a 90-day automated rerun of the orphan-candidate script and review unresolved candidates with their owners.

Require an accountable owner for every operational page you keep, and archive pages nobody will maintain.

Never Miss an Update

Fresh insights every week.

No spam. Unsubscribe anytime.

Your Thoughts

Share your thoughts.

Join the Discussion

Customise cookies