IT Runbook Template: How to Document Your Operations for Reliability
A runbook is an SOP for IT operations. The difference is context: runbooks are used under pressure — during incidents, after-hours, by engineers who may be unfamiliar with the specific system. The format requirements are different from a standard business SOP.
A runbook written for a calm afternoon review won't work for an on-call engineer at 2am. Here's how to write one that does.
Runbook vs. SOP: when to use which
An SOP covers repeatable operational procedures in normal conditions. A runbook covers:
- Incident response: What to do when an alert fires
- System maintenance: Scheduled tasks like database backups, certificate renewals, deployments
- Disaster recovery: Step-by-step procedures for system failures and data recovery
The key difference: runbooks are triggered by events or alerts and often used by engineers who didn't write them.
The 8 sections of a reliable runbook
1. Alert / Trigger condition
What triggers this runbook? Be specific: "PagerDuty alert: 'prod-api-latency-p95 > 2000ms for 5+ minutes'" or "Triggered monthly on the 1st by the ops-maintenance calendar."
This is the first thing the on-call engineer reads. If they can't immediately confirm this runbook matches their situation, they'll stop and search for a different one — wasting time.
2. Severity and initial assessment
What's the impact? Who should be notified immediately? What's the expected resolution time?
| Severity | User Impact | Initial Notify | Max Unresolved |
|---|---|---|---|
| P1 | Total outage | CTO, VP Eng | 1 hour |
| P2 | Degraded service | Engineering lead | 4 hours |
3. Prerequisites and access
What does the engineer need before starting? List every access requirement:
- AWS console access (ops-readonly role)
- VPN connected
- Datadog dashboard: [direct link]
- Slack channel: #incidents
Don't make them go find things. Every minute hunting for credentials is a minute not resolving the incident.
4. Diagnostic steps
Ordered investigation steps — not solution attempts. First understand what's actually wrong. Include:
- Commands to run (with exact syntax)
- Dashboards to check (with direct links)
- Logs to query (with sample queries)
- Expected normal values vs. abnormal values
Example:
Step 1: Check error rate
> kubectl logs -n prod deployment/api --since=15m | grep "ERROR" | wc -l
Normal: <10/min | Alert threshold: >100/min
5. Resolution procedures
Ordered resolution steps, with decision branches for the most common root causes. Use a clear "if/then" structure:
"If Step 4 shows memory > 90%: proceed to Step 5a (restart pod). If Step 4 shows error rate only on /checkout endpoint: proceed to Step 5b (rollback last deploy)."
6. Verification steps
How do you confirm the resolution worked? Don't end on "restart the service" — end on "confirm service is healthy by checking X, Y, Z."
7. Rollback procedure
If the resolution makes things worse, how do you undo it? Especially important for deployment runbooks.
8. Escalation contacts
If the runbook doesn't resolve the issue within X minutes, who do you call? Include name, role, and contact method (not just "escalate to the database team").
Common runbook mistakes
Writing for yourself. The person executing the runbook during an incident may not have written it. Assume they're competent but unfamiliar with this specific system.
No direct links. "Log into the monitoring dashboard" is not a runbook step. "Log into https://app.datadoghq.com/dash/1234 (bookmark in the ops-team Notion page)" is.
No commands. Describing what to do in prose is not useful for a technical runbook. Show the exact commands, including flags.
No expected outputs. Tell the engineer what a successful step looks like. "The command should return 'OK' if healthy."
Generate an IT runbook in 60 seconds
Incident response, system maintenance, and operational runbooks for IT teams.
Generate IT Runbook