Postmortem

Ingestion halt — 2026-07-30

Status: Draft · When: 2026-07-30 14:12–16:04 UTC · Owner: Platform

Summary

The single ingestion queue filled. Every producer sharing that path stopped accepting events for 1 hour 52 minutes. Warehouse freshness lagged. This is the fourth incident of this shape in six events.

Impact

Verified: pager timestamps; gateway 5xx from the 14:12–16:04 window.

Timeline (UTC)

Root cause

One queue, no failover. Backpressure had nowhere to go except the gateway. Scaling consumers recovered this instance; it does not remove the coupling. RFC 014 exists because of this chain.

Contributing factors

What went well

On-call found the lag dashboard on the first try. The runbook scale step worked.

Action items

  1. Ship RFC 014 (two queues). Owner: Platform. P1. Tracking: RFC 014.
  2. Page at 60s lag, not 120s. Owner: Reliability. P2.