Of the 21 third-party apps I audited in June and July 2026, at least 17 had no deploy gate: every push ships straight to production with nothing checking it first. When a release breaks, the fix depends on whoever remembers the steps. A runbook template turns deploy, rollback, key rotation, restore and each alert into a page anyone can follow.
What is a runbook, and what goes in a runbook template
A runbook is a written page that lets someone who did not build the system do one operational task under pressure: what triggers it, the exact steps, how to check it worked, how to undo it and who to call. A runbook template is that page with the fields fixed, so every runbook reads the same way.
Those 21 apps are the third-party apps I audited in June and July 2026: 11 third-party public vibe-coded apps audited exhaustively across all 12 pillars, plus a held-out set of 10 disjoint third-party apps the engine had never seen, audited blind. All 5 of my own apps had no deploy gate either. I chose these apps; they are not a random sample, so the count describes them and is not a rate for AI-built apps in general. Runbooks are one of the things a handover should leave behind, next to the rest of what a developer handoff looks like.
The fields below are my template. Every runbook on this page uses all ten, in this order, and the example column fills them in for a rollback.
| Field | What goes in it | Example line (rollback) |
|---|---|---|
| Title | The task, starting with a verb | Roll back a bad release |
| Trigger | The event that starts this runbook, and nothing else | Errors rise, or a customer reports a break, inside the watch window after a deploy |
| Owner and backup owner | Two named people, so one holiday never blocks the task | Owner: the lead engineer. Backup: the founder |
| Preconditions and access | Accounts, roles and tools needed before the first step | A role on the hosting account that can change production; access to the error tracker |
| Steps | Numbered, each with the exact command or click | Open the host’s deployments list, pick the last good release, confirm the rollback |
| The check | How you know it worked, in something you can see | The previous version is serving and the new error has stopped appearing |
| The undo | How to reverse this runbook if it made things worse | Promote the release you rolled back from, once it is fixed |
| Escalation | Who to call, and the condition that makes you call | Call the backup owner if the check still fails after the rollback |
| Last rehearsed | The date of the last full run on staging or production | The date of the last staging rollback |
| Evidence | A link to the log, the recording or the timed result | The staging run’s log and how long it took |
The documentation set for a vibe-coded app says each operations runbook should state the signal, business impact, first safe checks, containment options, decision owner, communication path, recovery, verification and escalation condition, and should be stored somewhere available when the app itself is down: that list is part of how to write technical documentation for a vibe-coded app. This template carries the same items under fewer headings: signal and impact sit in the trigger, first checks and containment in the steps, the decision owner in the owner field.
A runbook is not an SOP: an SOP sets the rule, and the runbook is the exact procedure for one system. Nor is it the automation kind. Azure Automation “supports several types of runbooks”, such as a “Textual runbook based on Windows PowerShell scripting”, and AWS Systems Manager automation runbooks define “the actions that Systems Manager performs on your managed instances and other AWS resources when an automation runs”. A machine runs those; a person runs the page this article is about.
In project management, a runbook is narrower than the project plan: the plan covers weeks of work, while the runbook covers one job on one live system, such as the cutover on launch day. In software development and DevOps, an application runbook is the same page with the app’s own triggers: a failed deploy, a queue that stops draining, an alert from the error tracker. Across tech, a good runbook document is short enough to read while something is on fire, and all runbook documentation earns its place the same way: someone else can follow it.
What goes wrong without it
Each of these starts as an ordinary week and turns into a bad one because the steps live in one person’s head.
| Situation | What happens with no runbook | The runbook that fixes it |
|---|---|---|
| The deploy only one person knows how to run | Every change waits on one calendar, and a holiday freezes releases | The deploy runbook |
| A bad release, and nobody is sure how to undo it | With no deploy gate, the release reached users unchecked, and the undo is improvised while they watch | The rollback runbook |
| An alert fires and nobody knows what to do with it | People guess, restart things, or mute it and hope | One runbook per alert |
| A leaked key, rotated by guesswork | The old key keeps working, or the new one breaks production mid-rotation | The key rotation runbook |
The second row is where the opener’s count bites: with no deploy gate, nothing stands between a push and production, so a release gate is worth adding before the rollback runbook is ever needed, and the release readiness checklist sets one out. The third row assumes an alert exists at all, and 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. That count comes from the same June and July audits of apps I picked rather than sampled, so it tells you what to check in your own app, not how common the gap is everywhere. Sending alerts somewhere a person will see them is the job of logging and monitoring.
In my reading, incident response for a small team of one or two people is these runbooks plus one agreed phone number and a status message, not a team structure with roles and rotations. The security version, with roles and the notification steps, belongs in an incident response plan template.
Can you give me an example of a runbook? The template, filled in five times
A runbook example for a small SaaS is five pages built from one template: deploy, rollback, key rotation, restore, and one page for each alert that can wake someone. Each page names its trigger, owner, steps, check, undo and escalation, and records the last time someone rehearsed it.
My working rule for how to create a runbook from the template is three steps. Do the operations task once for real and write the runbook as you go, every step exactly as you typed or clicked it. Then have someone who did not write it do the same task from the page alone while you watch and say nothing. Fix every step they had to ask about; a runbook prepared this way has already been run by someone other than its author.
The best runbook template is one your team will actually fill in, and this one is free to copy: a simple runbook template with ten fields and no download. An IT runbook template, an ops or service runbook, a software or production support runbook: each is this same template with a different trigger. A migration runbook template is the deploy runbook with the migration step as its own trigger.
The five blocks below assume a small SaaS on a managed host. Each one leaves the detailed procedure to the page that owns it.
Deploy runbook
| Field | Deploy runbook |
|---|---|
| Trigger | A merged pull request with green checks |
| Owner | The engineer who merged it; backup: the lead engineer |
| Steps | Promote to production through the release gate the team uses; run the migration step; run the smoke journeys; post the release note |
| The check | The version the app reports matches the release, and the smoke journeys pass |
| The undo | The rollback runbook |
| Last rehearsed and evidence | Every deploy is a rehearsal, so the evidence is the last deploy’s log |
A deployment runbook template needs no invented drill: the deploy itself is the rehearsal, which makes it the easiest runbook to keep true. The release gate in the first step is the one the release readiness checklist describes.
Rollback runbook
| Field | Rollback runbook |
|---|---|
| Trigger | Errors, or a customer report, inside the watch window after a deploy |
| Owner | Whoever deployed; backup: the lead engineer |
| Steps | The host’s own rollback, followed as the walkthrough below describes |
| The check | The previous version is serving and the error is gone |
| Escalation | The backup owner, if the check fails after the rollback |
| Last rehearsed and evidence | A timed rollback on staging, with the time written on this page |
The steps field stays one line on purpose. The walkthrough for Vercel and Netlify, the case where the host has no rollback button, and why a deployment rollback does not roll back the database are all in how to roll back a deployment. One point from there belongs in your preconditions field: on Vercel, which earlier deployments you can pick depends on the plan, and Hobby users can roll back to the immediately previous deployment. I’d rehearse the rollback on staging about once a month, timed, as my working rule, and write the cadence the team agrees on the page itself.
Key rotation runbook
| Field | Key rotation runbook |
|---|---|
| Trigger | A suspected leak, someone with access leaving, or the rotation schedule |
| Owner | The person who manages the secrets store; backup: the founder |
| Steps | Issue the new key; add it beside the old one; deploy; confirm the new key works in production; revoke the old key; confirm the old key is refused |
| The check | A request made with the old key is refused |
| Last rehearsed and evidence | The refused request for the last low-risk key rotated |
The order is the runbook; the detail per credential type, including the overlap window and what to do if the new key fails, is in how to rotate API keys safely. My working rule is to rotate one low-risk key about once a quarter so the page gets walked before it is needed for real.
Restore runbook
| Field | Restore runbook |
|---|---|
| Trigger | Data loss, corruption, or a failed migration |
| Owner | The person who manages the database; backup: the lead engineer |
| Steps | Choose the backup by time; restore into an isolated copy; check the restored data; decide how it reaches production; tell the people affected |
| The check | The rows and files you expect are present in the restored copy |
| Escalation | The founder, before anything is written back to production |
| Last rehearsed and evidence | The timed restore drill, with its time written on this page |
The steps field names the order and nothing more. The drill that proves a restore works, including what counts as a pass and how to check the restored data, is a restore drill for any backend. If the production database is gone, including the question of whether to restore over production, follow the first hour after deleting the production database. When the trigger is a failed migration, deciding between reversing it, fixing forward and restoring is the subject of database migration rollback, so this runbook never makes a restore the default. A disaster recovery (DR) runbook template is this restore page plus the disaster recovery checklist for SaaS.
One runbook per alert
| Field | Alert: error rate above your threshold for a few minutes |
|---|---|
| Trigger | The error-rate alert fires (the threshold is yours to set) |
| Owner | Whoever is on call; backup: the lead engineer |
| Steps | Open the error tracker; find the first new error; compare its start with the last deploy time; decide between rolling back (the rollback runbook) and fixing forward; update the status page |
| The check | The error rate is back under the threshold and the alert clears |
| Escalation | The backup owner, if the error has no link to a deploy |
| Last rehearsed and evidence | The last real alert handled from this page, or a test alert on staging |
My working rule is that every alert that can wake someone gets a page built from the alert runbook template; an alert with no page is either given one or turned off. Taken together, these pages are the on-call runbook, and in my view the runbook template is enough incident management tooling for a team of one or two to start with.
AWS’s post-event summary of the 2017 S3 disruption shows what one step can do with one wrong input. On the morning of February 28, 2017, the S3 team in US-EAST-1 was debugging an issue causing the S3 billing system to progress more slowly than expected. An authorized S3 team member “using an established playbook executed a command which was intended to remove a small number of servers” for one of the S3 subsystems used by the billing process. “One of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended”, and the servers removed by mistake supported two other S3 subsystems, the index and placement subsystems. AWS says the tool used “allowed too much capacity to be removed too quickly.” AWS says it modified the tool to remove capacity more slowly and added safeguards against taking any subsystem below its minimum required capacity level, and that it is auditing its other operational tools to ensure they have similar safety checks.
The lesson I take from it: a runbook step that can do damage needs a check before it runs, not only after. For a small SaaS, any step that deletes, restores over or revokes something should say what it will touch, and someone should read that back before it runs.
Where runbooks live: Confluence, GitHub, OneNote, SharePoint, Excel and Word
The tool matters less than the place. Keep the runbooks next to the code or in the wiki the team already opens every day, never in one person’s drive, and keep a copy readable when the app or its host is down, as the documentation article says.
To make a Confluence runbook template, a space administrator creates a page template from the space settings; Atlassian’s docs say “Only space administrators can create or edit templates in Confluence Cloud.” Those are Confluence’s page templates, and they can hold placeholder text that tells the writer what each field wants. On GitHub, the runbook template is a Markdown file in the repository’s docs folder, so each application’s runbooks change in the same pull request as its code, which is the DevOps habit worth copying. A OneNote runbook template is a page template, which Microsoft describes as “a page design that you can apply to new pages in your notebook”; you can also create your own, and Microsoft’s page covers OneNote for Microsoft 365, OneNote 2024, OneNote 2021 and OneNote 2016. For a SharePoint runbook template, I’d keep one finished runbook page as the model and copy it for each new runbook.
In Excel, the runbook template becomes one sheet per runbook with the ten field names down the first column; the deploy block above is a sample of one finished sheet. In Word, the same runbook template fields become headings. There is no runbook template to download here, free or otherwise: the ten-field table near the top of this page is the whole thing, and it pastes into Excel, a Word doc or a PPT slide as it stands, or prints to PDF.
The case for an SRE-style runbook template is made in one sentence of Google’s SRE book, which uses the word playbook: “When humans are necessary, we have found that thinking through and recording the best practices ahead of time in a ‘playbook’ produces roughly a 3x improvement in MTTR as compared to the strategy of ‘winging it.’” An AWS runbook template can mean this page for an app hosted on AWS, or one of the Systems Manager automation runbooks described above. The same runbook template serves ITIL change work or a project cutover, with the change request or the project milestone as its trigger.
The standing infrastructure list behind the runbooks
The standing infrastructure list is the one page every runbook leans on: each host and service with its account owner, each environment’s URL, each domain and DNS record, each scheduled job, each secret’s name and where it is stored, and each alert with where it routes.
This is the infrastructure documentation a small SaaS needs before anything longer. Every runbook points at this list instead of repeating it, which keeps an infrastructure runbook template short and means a changed DNS host is fixed in one place. The secrets line records each secret’s name and where it lives, never its value. The list’s picture is the architecture diagram, and this list is what you draw from when you work out how to create architecture diagrams for a web app. A server documentation template is this list with one row per server or service. An IT operations checklist template is this list plus the rehearsal cadences from the runbooks above, kept in one table so the next overdue rehearsal is easy to spot.
Security playbooks: the incident runbooks worth writing first
A playbook is a runbook for one kind of incident. A SOC runbook template is built for a security operations team watching many systems; a small SaaS gets further with three playbooks written from the incident response runbook template on this page. My working rule for the first three:
- A leaked secret: the key rotation runbook, plus who to tell and what to check in the logs for use of the old key.
- A suspected account takeover: revoke the account’s sessions, reset its credentials, then read the audit log for what the attacker touched.
- A data exposure: contain it, preserve the evidence, and start the notification clock the incident response plan template sets out.
Each of the three keeps the trigger, owner, check and rehearsal fields of any other incident runbook template. Treat the three as an incident response runbook checklist: each one exists, has an owner, and has a rehearsal date.
For the wider frame, NIST SP 800-61 is now Revision 3 (April 2025), whose incident response life cycle model is based on the six CSF 2.0 Functions: Govern, Identify, Protect, Detect, Respond and Recover. For cyber security playbook examples from a government, CISA publishes two, one for incident response and one for vulnerability response, written for federal civilian (FCEB) agencies; CISA says future iterations of these playbooks may be useful for organizations outside the FCEB.
How to verify it
A runbook is verified by walking it against the live configuration and by rehearsal: a staging rollback, a timed restore, a rotated key whose old value is refused, and a real deploy’s log. Someone who did not write the page runs it from the page alone. A runbook with no rehearsal date is a draft.
Five checks, each with the evidence worth keeping:
- 01 Walk every runbook against the live configuration: the host, the database, the secrets store and the alert rules. Fix each step that no longer matches, and keep a dated list of what changed
- 02 Link each runbook to its rehearsal evidence: the staging rollback time, the restore drill result, the refused request with the old key, a real deploy log. Keep those links on the page itself
- 03 Have someone who did not write the rollback runbook run it on staging from the page alone. Keep their time and every question they had to ask
- 04 Check that every last-rehearsed date sits inside its cadence. A page with no date is still a draft
- 05 Compare the standing list with the host account and project settings. Keep the date you checked it
The third check is the one that proves the page works without its author in the room. On the Production Hardening Sprint, deliverable 13.3 is verified this way: we walk through the runbooks against the delivered configuration and reference the rehearsal evidence.
Where the sprint does this
On the sprint, we document deployment, rollback, key rotation, backup restoration, and the response to each operational alert (deliverable 13.3). For the live technical handover, 13.4, we conduct and record a 60-minute walkthrough with the client team and technical advisors. The production readiness report, 13.1, accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Our included cover after handover is 14 calendar days of fixes for defects in the delivered work and 30 calendar days of async questions about the handover and architecture. Each of these is listed with its verify step in the published scope.
Runbooks are one part of a handover. The others are the items on the software project handover checklist, the runbook rows of the production readiness checklist, and a plan for how to run a technical handover session.
Common questions about runbooks
What are the different types of runbooks?
Runbooks sort into four practical types by what triggers them: routine ones such as deploy and key rotation, recovery ones such as rollback and restore, one per alert, and incident playbooks for security events. A separate meaning belongs to Azure Automation and AWS Systems Manager, where a runbook is a script or document the platform itself runs rather than a page a person follows.
What is a runbook vs sop?
A runbook says exactly how to do one task on one system, step by step, with the check and the undo; an SOP says what must happen and who is responsible for it, usually across a team or a process. In my reading, an SOP can name a runbook as the way to meet it, but nobody should need the SOP open to follow the runbook.
What is a runbook vs playbook?
A playbook covers one kind of incident and the decisions inside it, such as whether to take a service offline; a runbook covers one task. A playbook often calls runbooks: the leaked-secret playbook decides who to tell and when, then hands the actual rotation to the key rotation runbook. That split is my reading, not a standard.
What is a common use for runbooks?
A common use is the alert that fires outside working hours. The person who answers opens the page for that alert and follows it instead of guessing, so the fix does not depend on who happens to pick up the phone.
What is a devops runbook?
A DevOps runbook is a runbook kept with the code, changed in the same pull request as the thing it describes, and rehearsed as part of releases. That is my working rule for keeping it true: if a deploy changes a step, the same pull request changes the page.
Owning an app means being able to run it, change it and recover it without guessing. The sprint below leaves you with the runbooks and documentation to do that.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase