Engineering

OAuth Token Expiry: The Most Preventable Integration Outage

Credential expiry causes a disproportionate share of integration failures and is almost entirely avoidable. The failure modes, and the handling that removes them.

OAuth Token Expiry: The Most Preventable Integration Outage

Token expiry failures happen because refresh is treated as an error path rather than a routine operation. The fix is to refresh proactively before expiry, serialise refreshes across instances, persist rotated refresh tokens atomically, and alert on refresh failure rather than on the resulting 401s.

Key takeaways
  • Refresh before expiry on a schedule. Refreshing in response to a 401 is already too late.
  • Concurrent refreshes race. Serialise them, or a rotating provider will invalidate your token.
  • Alert on refresh failure. By the time 401s appear, the integration is already down.

Why this one is special

Most integration failures involve another party's behaviour. Token expiry is almost entirely self-inflicted, entirely predictable, the expiry time is handed to you in the token response, and still accounts for a large share of integration incidents. It is the clearest example of a failure that monitoring will not prevent and correct handling will.

The five failure modes

1. Reactive refresh

The integration refreshes only after receiving a 401. Every refresh is therefore preceded by at least one failed request, and under concurrency by hundreds. Some of those failures are user-visible.

Correct handling: record expires_in at issue time and refresh on a schedule at some fraction of the lifetime, commonly around three quarters, well before expiry.

2. Concurrent refresh races

Multiple workers detect expiry simultaneously and all call the refresh endpoint. With providers that rotate refresh tokens, the first call invalidates the token the others are holding, and the integration is now hard-down requiring manual reauthorisation.

Correct handling: serialise refresh with a distributed lock. One instance refreshes; the others wait and read the new token.

3. Lost rotated refresh tokens

Providers that issue a new refresh token on each use require you to persist it atomically. A crash between receiving the new token and committing it loses the credential permanently.

Correct handling: write the new refresh token durably before using the new access token, and treat that write as the commit point.

4. Silent scope loss

A refresh succeeds but returns a token with narrower scopes, because an administrator changed a policy or the app's permissions were revised. Calls now fail with 403 on some operations while others work, which looks like an intermittent provider fault.

Correct handling: assert the returned scopes match what you require on every refresh, and treat a mismatch as a failure.

5. User-bound credentials

The integration authenticates as an employee. They leave, the account is deprovisioned, the integration dies. Common, and entirely avoidable by using service accounts or app-level credentials where the provider supports them.

Alert on the right event

Most teams alert on 401 rates. By then the integration is already failing. Alert on refresh failure and on time until expiry falling below a threshold. Both are leading indicators and both are available before any request fails.

A correct refresh routine

  1. Store the access token, refresh token, absolute expiry timestamp and granted scopes together.
  2. A scheduled task refreshes any credential past roughly three quarters of its lifetime.
  3. Refresh happens under a lock keyed on the integration, so only one instance refreshes.
  4. On success, verify scopes, then persist the new refresh token atomically before use.
  5. On failure, alert immediately with the integration name and the provider's error: this is an incident, not a retryable blip.
  6. Keep the 401 path as a fallback only, and record it when it fires, because it firing means the scheduled refresh did not work.

Why it keeps happening

Because it works in development, where one process holds one token and nothing rotates. The failure modes above all require concurrency, time, or an administrator: none of which are present when the integration is written. The discipline is to treat credential lifecycle as part of the integration's design rather than as error handling, and to record each integration's token lifetime and rotation behaviour in the inventory alongside its other operational facts.

Rotation behaviour varies, so record it

Providers differ in ways that determine whether your refresh logic is correct, and the differences are rarely prominent in the documentation:

  • Does the refresh token rotate on use? If so, persistence must be atomic and must happen before the new access token is used.
  • Is there a grace period? Some providers honour the previous refresh token briefly, which makes a concurrency race survivable. Most do not.
  • Does the refresh token itself expire? An absolute lifetime means the integration needs periodic re-authorisation regardless of activity, which should be a calendar item rather than a surprise.
  • Is there an idle expiry? Integrations used infrequently can die from disuse.
  • Are concurrent refreshes rejected or tolerated? Determines whether a lock is mandatory or merely advisable.

Five lines per provider in the inventory. The question that catches most teams is the third, because an integration that has worked for a year can stop on a date nobody recorded.

Frequently asked questions

Why do OAuth integrations fail after working for months?

Usually because refresh is handled reactively or because a refresh token rotated and was not persisted. Both work fine until a concurrency race, a crash at the wrong moment, or an administrator policy change exposes them: none of which occur during development.

How do you prevent token expiry outages?

Refresh proactively on a schedule before expiry rather than in response to 401s, serialise refreshes with a lock so concurrent workers do not race, persist rotated refresh tokens atomically before use, verify returned scopes, and alert on refresh failure rather than on resulting errors.

Should integrations authenticate as a user account?

No, where the provider offers an alternative. User-bound credentials break when that person's account is deprovisioned. Service accounts or app-level credentials remove an entire class of failure at effectively no cost.

Stop rediscovering the same integration failure

Traxivo correlates the signals your tools already produce into one incident timeline, recognises a recurrence as a recurrence, and drafts the follow-up with the evidence attached. Nothing is sent without a named approver.

See how Traxivo works Browse use cases

Related reading