Discovery

Should we split ingestion, or is the problem somewhere else?

Decision this enables: go to RFC 014 / stop / reframe · Window: 2026-07-01 – 2026-08-08 · Owner: Platform

Goal of this discovery

Decide whether a queue split is the right next investment, or whether detection, consumer capacity, or producer behaviour would remove more halt time cheaper.

Problem as reframed

Not "we need two queues". The problem: when ingestion backpressures, the whole platform stops accepting events, and that has happened four times in six incidents.

Users and context

Operators (Platform on-call). Producer teams who feel the 503s. Downstream analysts who wait on the warehouse. Wider journey: checkout write → event → warehouse → inventory views.

Evidence

Opportunities

Constraints

Alternatives to building

Page earlier and scale consumers — cheaper, does not stop a full fill. Doing nothing remains the default if this discovery recommends stop.

Assumptions still open

  1. Two queues actually fail independently — untested. Highest build risk.
  2. Shed-and-alert is acceptable to Checkout. Not yet asked.

Recommendation

Proceed to RFC 014, and also lower the lag alert in parallel. Stopping would accept another halt this quarter; the cost of being wrong on the split is one extra queue to run, not a new user-facing surface.

What would be tested next

Shadow-consume on a second queue for a week (RFC 014 rollout). Ask Checkout whether 202-shed is tolerable.