Back to Blog
Best Practices

One Tenant's Webhook Flood Shouldn't Become Everyone's Outage

One Shopify store sent 30 million events in five minutes, and every other store's webhooks waited hours behind it. Here is why a single shared delivery queue turns one noisy tenant into a platform outage, and how per-tenant fairness keeps everyone else moving.

WebhookVault Team · Webhook Infrastructure Experts10 min read
Close-up of black Ethernet cables plugged into the ports of a silver network device, with small amber port lights glowing against a dark background

Thirty million events from one store

One store sent 30 million events in five minutes. That is the incident Hookdeck described when it introduced Delivery Groups: a customer running webhook delivery for hundreds of Shopify stores watched a single shop flood the pipe. Every other store's order, fulfillment, and inventory webhooks sat in the queue for hours.

Nothing was down. Receiver and broker were both fine. The workers were busy doing exactly what they were told: draining the queue in arrival order. That is the problem. A FIFO queue treats every message the same, and your customers aren't messages. They're tenants.

If your webhook pipeline has a single queue per destination, or a single queue full stop, you already have this bug. You just haven't met the tenant who triggers it.

A single queue is a single line

Picture the only checkout lane in a shop. One customer shows up with four thousand items, and the person behind them holding a single bottle of water waits anyway. Open more lanes and the big cart gets split across all of them. Everyone is still waiting for the big cart.

Horizontal scaling does the same to a shared webhook queue. More workers drain the backlog faster, in arrival order. If one tenant put 99% of the backlog there, 99% of your new capacity goes to that tenant. The small tenant's single orders/create event is still number thirty million and one in line.

The cruel part is how the metrics look. Throughput hits a record, errors sit at zero, and queue depth is large but dropping. Every graph you built for webhook health says the system is coping. Meanwhile, hundreds of merchants are watching their inventory drift and opening support tickets.

Two kinds of noisy neighbor

There are two distinct ways one party can starve everyone else, and they need different fixes.

The first is the loud tenant. One customer produces far more events than the rest: a bulk import, a migration script, a sync loop that updates the same product forever. Volume is the weapon. Every event is cheap to deliver, but there are millions of them in front of yours.

The second is the slow endpoint, and it only matters if you send webhooks to many destinations. One customer's receiver starts taking 25 seconds to answer, or stops answering and lets every request hit your timeout. Volume is normal. It eats worker time. A worker holding an open connection to a dead endpoint delivers nothing to anyone else.

Receivers mostly meet the first kind. Senders meet both, and the slow endpoint is more common. A loud tenant shows up in your event counts. A slow endpoint shows up as a worker pool that is somehow always full.

Isolate destinations before you isolate tenants

If you send webhooks, start with the destination. Put a concurrency cap on every endpoint so no single URL can hold more than a handful of your workers at once.

The arithmetic is unforgiving. Say you run 200 delivery workers with a 30 second timeout, and 40 customer endpoints start hanging. Without a cap, those 40 endpoints slowly absorb the pool. Each hung request pins a worker for the full 30 seconds, retries pile onto the same dead hosts, and within minutes the healthy endpoints are fighting over whatever workers are left. With a cap of four in-flight requests per endpoint, the dead endpoints can hold at most 160 workers. That is still bad. Cap at two and they hold 80, and the remaining 120 keep everyone else on schedule.

const MAX_IN_FLIGHT_PER_ENDPOINT = 2

async function tryDeliver(job: DeliveryJob): Promise<'sent' | 'deferred'> {
  const key = `inflight:${job.endpointId}`
  const current = await redis.incr(key)
  await redis.expire(key, 60)

  if (current > MAX_IN_FLIGHT_PER_ENDPOINT) {
    await redis.decr(key)
    return 'deferred' // requeue with a short delay, do not count as an attempt
  }

  try {
    await deliver(job)
    return 'sent'
  } finally {
    await redis.decr(key)
  }
}

The deferral doesn't count as a failed attempt. A delivery that never left your building shouldn't burn one of the customer's retries. Count it, and a hung endpoint next door ends up getting your healthy customers' events dead-lettered.

For this case, reach for a concurrency cap before a rate limit. A limit of 10 requests per second does nothing about an endpoint that holds each request for 30 seconds. A cap bounds the thing you actually run out of, which is workers. Our rate limiting guide covers the token-bucket side. This is the half the bucket doesn't touch.

Fairness inside one destination

Destination isolation does nothing for the Hookdeck case. There, every event goes to the same destination: one platform receiving webhooks on behalf of hundreds of stores. The noise and the victims share a URL.

The fix is to split one destination's traffic by a key inside the event. Hookdeck's Delivery Groups do precisely this: you point at a field in the body, headers, query, or path, something like body.shop_id, and every tenant gets its own delivery rate. When events arrive faster than the destination can absorb them, the gateway rotates between groups, so a spike from one tenant doesn't starve the rest. Idle groups don't reserve capacity, which matters when most of your tenants are quiet most of the time.

You don't need a vendor for this. You do need to stop treating the backlog as one thing. It's a set of per-tenant backlogs that happen to share workers.

Round-robin across tenant lanes

The simplest version that works: one queue per tenant, plus a ring of tenants that currently have work. Workers pop the next tenant from the ring, take a small batch from that tenant's queue, and put the tenant back at the end of the ring if anything is left.

async function nextBatch(batchSize = 10): Promise<Event[]> {
  const tenant = await redis.lpop('active-tenants')
  if (!tenant) return []

  const raw = await redis.lpop(`tenant:${tenant}:events`, batchSize)
  const events = (raw ?? []).map((r) => JSON.parse(r) as Event)

  const remaining = await redis.llen(`tenant:${tenant}:events`)
  if (remaining > 0) await redis.rpush('active-tenants', tenant)

  return events
}

async function enqueue(event: Event) {
  const tenant = event.shopId
  const len = await redis.rpush(`tenant:${tenant}:events`, JSON.stringify(event))
  if (len === 1) await redis.rpush('active-tenants', tenant)
}

With this in place, the store that sent 30 million events gets one turn per rotation, like everybody else. The store with one event waits for at most one lap of the ring. Total throughput is identical. The distribution of latency is completely different.

Two details decide whether it holds up. The check-then-push in enqueue races under concurrency, so wrap it in a Lua script or a MULTI block, or a tenant will sometimes be listed twice or not at all. And batch size is your fairness knob. Bigger batches help the loud tenant's throughput and lengthen everyone else's lap. Start small.

Paid tiers? Use weighted round-robin: two slots in the ring, or twice the batch size. Keep the weights coarse. Fine-grained priority is how these systems drift back into a single line.

Picking the fairness key

The key you split on decides who is protected from whom, and getting it wrong is easy.

Split by shop, account, or workspace: the unit your customers recognise as "mine." Splitting by event type protects orders/create from products/update. Nice. The loud tenant's orders still bury the quiet tenant's orders. Don't split by something with unbounded cardinality either, like a resource id. A million lanes with one event each is a FIFO queue with extra steps and a lot of Redis keys.

Make sure the key is present on every event, including the ones you didn't think about. The events with a missing key all land in one shared fallback lane, and that lane becomes the new noisy neighbor. Count how many events hit it. If that number climbs, someone added an event type without the field.

On the receiving side, you often can't trust the key until you've verified the signature. Do verification at ingest, then derive the tenant, then enqueue. A forged request shouldn't be able to choose which lane it jumps into.

Retries belong in a separate lane

Retries are their own kind of noisy neighbor. When an endpoint fails, every failed event goes back into a queue, and if that queue is the same one fresh events use, a broken endpoint's retry traffic competes with first attempts for everyone else.

Keep fresh deliveries and retries in separate lanes, and give the fresh lane priority. A first attempt is more likely to succeed than a fifth try against an endpoint that has failed four times. Spend capacity on the likely winner.

This also tames the retry storm when a large endpoint recovers. Without a separate lane, the moment it comes back, hours of deferred retries hit it in one wave. With a separate lane under its own concurrency cap, the backlog drains at a pace the endpoint can survive, which keeps it from tipping straight back over and ending up disabled by the sender.

Watch backlog age per tenant

Queue depth is the wrong number to alert on when you have tenants. If fairness works, a depth of 30 million can be one tenant's bulk import that nobody else notices. A depth of 300 can be a disaster if all 300 belong to your largest customer and the oldest is six hours old.

The metric that matches what customers feel is the age of the oldest undelivered event, per tenant. Emit it as a gauge with the tenant as a label, then alert on the maximum across tenants and on how many tenants exceed a threshold.

for (const tenant of await redis.lrange('active-tenants', 0, -1)) {
  const head = await redis.lindex(`tenant:${tenant}:events`, 0)
  if (!head) continue
  const ageSeconds = (Date.now() - JSON.parse(head).receivedAt) / 1000
  backlogAge.set({ tenant }, ageSeconds)
}

Mind the cardinality. A tenant label is fine for a few hundred tenants and painful for a hundred thousand. At that scale, report the top N oldest tenants plus a histogram of ages.

The two alerts tell you different things. One tenant with a large age means that tenant is noisy or its endpoint is broken, which is their problem and maybe a support conversation. Many tenants with growing ages at once means fairness is failing or total capacity is short, which is yours.

What fairness won't fix

Per-tenant fairness moves the waiting around. If total capacity is below total demand, everyone still waits, just more evenly. Fairness buys you time to scale. You still have to scale.

It also touches ordering. Inside a tenant's lane, arrival order holds. Across tenants it doesn't, and if you relied on a global order, you lost it the day you started a second worker. If ordering matters per resource, the out-of-order delivery problem is still yours to solve with version checks in the handler.

And to the loud tenant, fairness looks like you broke something. Their bulk import used to finish first. Now it takes hours. That's correct behavior, so tell them. A tenant whose events get paced on purpose, with nothing on a dashboard saying why, will open an incident. Give them a backlog-age view of their own and the same pacing becomes something they expected.

Frequently asked questions

Isn't rate limiting each tenant enough on its own? A per-tenant rate limit caps how fast one tenant drains, but it does nothing to reorder a backlog that already exists. If the loud tenant's events are already queued ahead of yours in a shared queue, they still go first, just more slowly. You need separate lanes for the limit to protect anyone.

Should the sender or the receiver handle fairness? Both, for different neighbors. The sender protects its other customers from one slow or broken endpoint with per-destination concurrency caps. The receiver, or the gateway in front of it, protects its own tenants from each other, because the sender sees one destination and has no idea there are hundreds of stores behind it.

How many tenant lanes can Redis handle? Lanes are just keys, so tens of thousands of active lists are routine. The cost is in operations per event, not in the number of lanes. Empty lanes cost nothing because a tenant drops out of the ring as soon as its list is empty.

What happens to a webhook that is missing the tenant key? It needs a defined home, usually a shared fallback lane with its own alert. Rejecting it outright risks dropping legitimate events, and silently mixing it into a random tenant's lane hides the bug. Count fallback hits and treat a rising count as a schema problem upstream.

Related posts