Practice

Integration Runbooks: What Actually Belongs in Them

Most runbooks are written once and never used. What makes an integration runbook worth opening at 2am, and the five sections that earn their place.

Integration Runbooks: What Actually Belongs in Them

An integration runbook is useful only if it is written for someone with no context under time pressure. Five sections earn their place: how to confirm the failure is real, how to establish scope, what to do immediately, who to contact with what, and the known history of this integration.

Key takeaways
  • Write for a stranger at 2am, not for the person who built it.
  • Lead with confirmation, the most common error is responding to a failure that is not happening.
  • Known history is the most valuable section and the one almost always missing.

Why most runbooks fail

They are written by the engineer who built the integration, immediately after building it, while holding all the context. The result reads as a reminder rather than as instructions, and assumes knowledge the reader does not have. It is then not opened for eighteen months, by which point three of its steps reference systems that no longer exist.

The test is simple: could someone who has never touched this integration follow it, at 2am, without calling anyone? Most runbooks fail on the first paragraph.

Five sections

1. Confirm it is real

Start here, because the most common error in integration response is acting on a failure that is not occurring: a stale dashboard, a monitoring artefact, a report from a user who was looking at cached data.

Give one concrete check that takes under two minutes and produces an unambiguous answer: a specific query, a specific request to run, a specific place to look. Include what a healthy result looks like, because the reader does not know.

2. Establish scope

Which operations, which records, which customers, and since when. Give the exact query or filter that answers each. Scope determines everything downstream (whether to page anyone, whether to notify customers, whether this can wait until morning) and it is frequently skipped in favour of jumping to a fix.

3. Immediate actions

What to do now, in order, with explicit stopping conditions. Distinguish clearly between actions that are safe to take unilaterally and actions that need authorisation. If disabling the integration is an option, say so and say what breaks if you do.

Include how to verify each action worked. An instruction without a verification step leaves the reader unsure whether to continue.

4. Who to contact, and with what

The internal owner. The vendor's contractual support channel, not the one that is easiest to reach. What the vendor will require before escalating, which is covered in how to escalate to a vendor. Having that list in the runbook means the first message is complete rather than the first of four.

5. Known history

The section almost nobody writes and the one that most often resolves the incident. Previous occurrences, what caused them, what fixed them, and, critically, any workaround still in place that might be distorting the current symptoms.

Why history belongs in the runbook

A large share of integration incidents are recurrences. If the reader can match current symptoms to a prior occurrence in the first five minutes, the remaining four sections are often unnecessary. This is also why history should be generated from incident timelines rather than typed by hand: hand-maintained history is the first thing to go stale.

What to leave out

  • Architecture explanation. Link to it. Nobody reads a design document during an incident.
  • Every possible failure mode. Cover the three or four that actually happen. Exhaustive runbooks are not read.
  • Anything that duplicates a dashboard. Link to the dashboard.
  • Credentials. Reference where they live; never inline them.

Keeping them honest

Runbooks decay silently. Two cheap practices prevent most of it: after every incident, the responder updates the runbook as part of closing out, particularly the history section, and once a quarter someone who did not write it attempts to follow it end to end. The second practice finds broken links and missing context that the author cannot see, and it takes twenty minutes.

Testing the runbook

Runbooks decay silently, and the author is structurally unable to notice because they supply the missing context from memory. The only reliable test is to have someone else follow it.

Once a quarter, pick one runbook and ask an engineer who has never touched that integration to work through it end to end against a non-production environment, while the author watches without speaking. The instruction not to speak is the important part.

What this reliably surfaces:

  • Links to dashboards, queues or consoles that no longer exist
  • Steps that assume access the reader does not have
  • Instructions with no verification step, leaving the reader unsure whether to continue
  • Missing context around what a healthy result actually looks like

Twenty minutes per quarter per runbook. The alternative is discovering the same gaps during an incident, with a customer waiting and the author asleep.

Frequently asked questions

What should an integration runbook contain?

Five sections: a fast unambiguous check that the failure is real, how to establish scope and start time, immediate actions with verification steps and authorisation boundaries, who to contact and what they will require, and the known history of previous occurrences including any workaround still in place.

Why do runbooks stop being useful?

They are written by the person with full context and read as reminders rather than instructions, then decay as referenced systems change. Having someone who did not write the runbook follow it end to end once a quarter surfaces the gaps the author cannot see.

Should runbooks list every possible failure mode?

No. Exhaustive runbooks are not read under time pressure. Cover the three or four failure modes that actually occur for that integration and link out to deeper material for anything rarer.

Stop rediscovering the same integration failure

Traxivo correlates the signals your tools already produce into one incident timeline, recognises a recurrence as a recurrence, and drafts the follow-up with the evidence attached. Nothing is sent without a named approver.

See how Traxivo works Browse use cases

Related reading