- Integration
- A connection between two software systems that are owned or operated separately. The defining characteristic is that a contract exists across an ownership boundary, which is what makes failure harder to diagnose than failure inside one system.
- Integration observability
- The ability to explain why a connection between two systems failed using evidence you already collect, including evidence originating on the other side of the boundary.
- Detection gap
- The interval between a failure's first occurrence and the moment a human first understood it was happening. For integrations this is routinely the largest component of total incident duration.
- Silent failure
- A failure that produces no error signal. The process reports success while producing incomplete or incorrect results: a pipeline run that moves a tenth of the expected rows, or a webhook endpoint that acknowledges an event it never processed.
- Partial failure
- A failure affecting a subset of operations, records or customers. Rarely crosses an alerting threshold, which is why partial failures typically run far longer than total ones.
- Silent 2xx
- A webhook endpoint returning a success status before the event is durably stored. The sender records delivery as successful; any later processing failure is invisible to both parties.
- Retry storm
- Amplification of load against an already-degraded dependency caused by retry logic, converting partial failure into total failure. Worsened when many clients retry in synchrony.
- Jitter
- Randomisation added to a retry backoff interval so that clients which failed simultaneously do not retry simultaneously. Its absence is the most common defect in otherwise correct retry logic.
- Idempotency key
- A client-generated identifier submitted with a request so the provider can recognise a duplicate. Must be stable across every retry of the same logical operation; regenerating it per attempt defeats the mechanism.
- Circuit breaker
- A control that stops calls to a dependency after a threshold of consecutive failures, waits a cooling period, then probes with a single request. Converts slow cascading failure into fast contained failure.
- Schema drift
- A change in the structure of a provider's responses: a field changing type, becoming optional, or silently returning null. Breaks downstream logic without producing any HTTP error.
- Replay window
- How far back a provider's API allows events to be re-requested after a delivery outage. Determines whether an outage is recoverable, and is best established before an incident rather than during one.
- Reconciliation
- Periodic comparison of record counts or contents across an integration boundary to detect divergence. The only method that reliably catches silent failure.
- Watermark
- The high-water mark a pipeline uses to track what it has already processed. A watermark that fails to advance produces successful runs that process nothing, often for weeks.
- Backfill
- Reprocessing of historical data to repair a gap left by a failure. Depends on the provider's replay window and on rate limits that frequently make large backfills impractical.
- Rate limit
- A provider-imposed cap on request volume, normally signalled by a 429 status and a Retry-After header. Under growth, a limit that was never close becomes a source of partial failure.
- Refresh token rotation
- A provider issuing a new refresh token each time one is used. Requires the new token to be persisted atomically before the access token is used; a crash in between loses the credential permanently.
- Scope loss
- A credential refresh succeeding but returning narrower permissions than before, usually after an administrator policy change. Produces errors on some operations while others continue to work, which resembles an intermittent provider fault.
- Integration inventory
- A maintained list of every connection between your systems and anything outside them, recording owner, criticality, authentication, failure mode, evidence location and recovery path.
- Incident timeline
- An ordered record of every signal relating to one integration fault, drawn from every system that observed it, including exchanges with the counterparty and any workaround applied.
- Recurrence
- A new occurrence matching the signature of a previous incident. Warrants a different response from a first occurrence, and is the metric most worth driving down.
- Mean time to detect (MTTD)
- Average interval from first occurrence to human awareness. For integrations this is usually larger than mean time to repair and is the more tractable target.
- Mean time to resolve (MTTR)
- Average interval from first occurrence to verified resolution. Measuring from first report instead of first occurrence discards the detection gap and produces a flattering, unusable number.
- Error budget
- An agreed allowance of failure over a period. For integrations it is only meaningful if its definition of failure includes partial and silent failure, not just unavailability.
- Blast radius
- The scope of a failure in operations, records and customers affected. Drives vendor prioritisation more reliably than severity language does.
- Breaking-change notice
- A contractual commitment to warn a named technical contact before changing schemas, authentication or rate limits. Often the most valuable clause available to a smaller buyer, and among the cheapest for a vendor to grant.
Put this into practice without the manual overhead
Traxivo keeps the inventory, the timeline and the vendor history current as a by-product of handling the signals your tools already produce.
See how Traxivo works Browse use cases