Practice
Alert Fatigue in Integration Monitoring
Adding a threshold per metric produces more alerts than any small team can process, and within a quarter they are all ignored. What to alert on instead.
Alert fatigue is the predictable result of alerting per metric rather than per conclusion. The remedy is to raise an alert only when correlated evidence supports a specific action, and to route everything else to a review surface that is read deliberately rather than interrupted by.
- An alert should assert a conclusion, not report a measurement.
- If nobody would act at 3am, it is not an alert. It is a record.
- Correlation reduces alert volume. Adding monitors increases it.
How teams get here
Nobody sets out to create four hundred alert rules. Each one is added after an incident, by a reasonable person, to catch that specific failure next time. The rules accumulate, each firing occasionally, until the aggregate volume exceeds what the team can triage. At that point the alerts stop being read, not through negligence, but because reading them is no longer possible.
The damage is not the wasted attention. It is that the next real signal arrives into a channel that everyone has learned to ignore.
The structural mistake
Most integration alerting attaches a threshold to a metric: error rate above a value, latency above a value, queue depth above a value. Each rule sees one signal from one system, which means it cannot distinguish a meaningful event from noise, so it is tuned either to fire too often or to miss things.
A tail latency increase on its own is usually nothing. A tail latency increase, plus a change in the response schema hash, plus a rise in 422s, plus a support ticket mentioning the same integration is a conclusion: the provider shipped a change. Only the conclusion deserves to interrupt someone.
Three tiers
Classify every signal into one of three destinations, and be strict:
Interrupt
Something is broken, a human must act now, and the action is known. Customer-facing failure, or a credential failure that will become customer-facing within the hour. These should be rare enough that each one is read.
Review
Something changed and a human should look within a day or two. Recurrence of a known signature, a schema change, quota consumption trending toward a limit. This is where the large majority of integration signals belong, and most teams have no such surface, which is precisely why these end up as interrupts.
Record
Evidence worth keeping for the timeline but not worth anyone's attention now. Individual retries, transient 5xx, single failed deliveries that later succeeded.
For every existing alert rule, ask whether you would want to be woken by it. If not, it is not an interrupt: move it to review. Most teams find that the large majority of their rules fail this test, and moving them is the single fastest way to make the interrupt channel readable again.
What to alert on for integrations specifically
- Absence of expected volume: the highest-value integration alert, and one threshold-based monitoring almost never provides.
- Credential refresh failure: a leading indicator with a known action.
- Recurrence of a known signature: materially different from a first occurrence and deserving a different response.
- Reconciliation mismatch: conclusive evidence of data loss.
Notice that none of these is a threshold on an instantaneous metric. Each requires either a baseline or memory of previous events.
Why this is hard without help
Correlating signals across monitoring, ticketing and mail, remembering signatures across months, and judging when accumulated evidence crosses into action is work that does not fit in a threshold rule and does not fit in a small team's spare capacity either.
It is the work Traxivo does: watching the record the team already produces, recognising when scattered signals describe one fault, and raising it once, with the history attached, rather than emitting an alert per signal. The measure of success is that it surfaces fewer things, not more.
Auditing the rules you already have
A practical exercise that takes an afternoon and usually removes most of the noise. Export every alert rule and, for each, record three things: how many times it fired in the last ninety days, how many of those led to an action, and what that action was.
The output sorts itself into four groups:
- Fired often, never actioned. Delete or move to review. This is normally the largest group by a wide margin.
- Fired often, always the same action. Automate the action, or raise the threshold until the alert means something different.
- Never fired. Either the failure never occurred or the rule is broken. Test it. Rules that have never fired are frequently misconfigured and provide false assurance.
- Fired rarely, always actioned. Keep. These are the real ones, and in most estates there are surprisingly few.
Teams running this for the first time typically find the fourth group is under a tenth of their rules. That ratio is the problem stated numerically.
Frequently asked questions
What causes alert fatigue?
Alerting per metric rather than per conclusion. Each rule observes one signal from one system and cannot distinguish meaningful events from noise, so rules are tuned to fire too often. The volume eventually exceeds what the team can triage and the channel stops being read.
What should trigger an interrupt-level alert for an integration?
Customer-facing failure, or a leading indicator with a known action and a short fuse such as credential refresh failure. Everything else (schema changes, recurrence of known signatures, quota trends) belongs on a review surface examined deliberately within a day or two.
How do you reduce alert volume without missing failures?
Correlate signals into conclusions before alerting, and alert on absence of expected activity rather than on error thresholds. Correlation reduces volume while improving coverage, whereas adding monitors increases volume and degrades attention.
Stop rediscovering the same integration failure
Traxivo correlates the signals your tools already produce into one incident timeline, recognises a recurrence as a recurrence, and drafts the follow-up with the evidence attached. Nothing is sent without a named approver.
See how Traxivo works Browse use cases

