Metrics
MTTR for Integration Incidents: How to Measure and Reduce It
Standard MTTR understates integration incidents because the clock starts when someone notices. Split the measurement into detect, diagnose, escalate and resolve to see where the time actually goes.
Mean time to resolve for integration incidents should be measured from first occurrence, not first report. Measured that way it is typically several times larger than teams expect, and the largest component is almost never engineering time: it is detection delay and waiting on a counterparty.
- Start the clock at first occurrence. Starting at first report hides the largest component.
- Split MTTR into detect, diagnose, escalate and resolve. The bottleneck is rarely where teams assume.
- Time spent waiting on a vendor is still your MTTR. It is also the most reducible portion.
The measurement problem
Mean time to resolve is conventionally measured from the moment an incident is declared. For incidents inside your own boundary that is defensible: your monitoring declares them within seconds. For integrations it quietly discards the most expensive phase.
A sync that has been dropping records for eleven days before anyone notices, then is fixed in two hours, records an MTTR of two hours. The number is accurate and useless.
A four-part split
Measure four intervals rather than one:
- Time to detect. First occurrence to human awareness.
- Time to diagnose. Awareness to a correct understanding of cause and scope.
- Time to escalate. Diagnosis to the counterparty acknowledging it, where the fault is theirs.
- Time to resolve. Acknowledgement to the fix being live and verified.
Collected over a quarter, this split almost always shows the same shape: detection and escalation dominate, and engineering time, the part teams instinctively optimise, is a minority of the total.
Reducing each component
Detection
This is the largest and the most tractable. The levers are correlation across tools and alerting on absence rather than on errors: covered in why integration failures go undetected. Adding more thresholds makes this worse, not better.
Diagnosis
Diagnosis time collapses when a recurrence is recognised as a recurrence. The second occurrence of a known fault should take minutes, and it takes days when the history lives in the memory of whoever handled it last. This is the direct payoff of maintaining an incident timeline.
Escalation
The slowest phase for most SMB and midmarket teams, because the counterparty has no contractual reason to prioritise you. What moves it: a first message that contains everything needed to reproduce, addressed to the right queue, referencing prior occurrences. What does not move it: following up more often. See how to escalate to a vendor.
Resolution
Mostly outside your control once the fault is theirs. The one thing within your control is verification: confirming the fix holds rather than closing on the vendor's assurance. A substantial share of reopened integration incidents are fixes that were never verified.
Track how many incidents in a quarter are recurrences of a previous signature. If that fraction is not falling, MTTR improvements are cosmetic: you are getting faster at handling the same fault repeatedly rather than eliminating it.
What good looks like
Absolute targets are not portable across teams; the distribution of integrations matters too much. The useful comparisons are internal:
- Detection time trending down quarter over quarter.
- Second-and-subsequent occurrences resolving substantially faster than first occurrences.
- Recurrence rate falling.
- Escalation time distinguishable per vendor, which tells you which relationships need renegotiating rather than more patience.
That last one is the measurement most teams never make, and it is the one that converts an engineering metric into a commercial argument. Traxivo records these intervals as a by-product of assembling the timeline, which is the only way they get captured consistently in teams without a dedicated incident function.
Instrumenting the four intervals
Each interval needs a different source, which is why so few teams capture all four:
- First occurrence comes from raw telemetry, searched backwards once the fault is understood. It cannot be captured prospectively, only recovered, which is why log retention quietly determines whether this metric is available at all.
- Detection is the timestamp of the first human action: a ticket, a message, a page. Usually the easiest of the four.
- Escalation is the counterparty's first substantive reply, not your first message. The distinction matters: the gap between them is the part you can influence with better evidence.
- Resolution is verification against your own telemetry, not the vendor's claim. Closing on assurance is how incidents get reopened.
Collect them per incident for a single quarter before trying to improve anything. The shape of the distribution tells you where to spend, and it is rarely where the team expects.
Frequently asked questions
How should MTTR be measured for integration incidents?
From first observed occurrence to verified resolution, split into time to detect, diagnose, escalate and resolve. Measuring from first report rather than first occurrence discards the detection gap, which is usually the largest single component.
What is a good MTTR for integration failures?
Absolute targets do not transfer well between teams because the mix of integrations differs too much. Internal trends are more useful: falling detection time, recurrences resolving faster than first occurrences, and a falling recurrence rate.
Should time waiting on a vendor count toward MTTR?
Yes. It is time the failure is unresolved and customers are affected. Excluding it produces a flattering number and removes the pressure to improve the escalation path, which is often the most reducible part of the total.
Stop rediscovering the same integration failure
Traxivo correlates the signals your tools already produce into one incident timeline, recognises a recurrence as a recurrence, and drafts the follow-up with the evidence attached. Nothing is sent without a named approver.
See how Traxivo works Browse use cases

