webhook-proxy

A middleware layer between our services and Discord webhooks. It queues events per target, respects 429 limits and retries 5xx errors, so notifications do not get lost during traffic spikes. A single point through which all outbound communication passes.

webhook-proxy
TL;DR

A single egress point for all outbound communication to Discord. It queues events per target, respects 429 limits with proper backoff and retries 5xx errors instead of dropping notifications during a traffic spike. The upshot is easy to state: an event does not vanish just because a rate limit happened to hit.

Overview

Our services say a lot on Discord: monitoring alarms, event notifications, operational logs. Each of them goes out over a webhook, and Discord webhooks have hard rate limits. As long as traffic is calm, nobody feels it. The problem starts exactly when the notification is needed most - during a spike, when several services fire at once.

webhook-proxy came about to stop gluing limit handling into every service separately. Instead, all outbound communication goes through one point that takes delivery on itself: it queues, waits when it must, retries when it is worth it and does not lose the event along the way. For the services behind it nothing changes - they send as before.

The problem you only see under load

Under calm traffic naive webhook sending works flawlessly, and that is exactly why it is treacherous. The code fires a request, gets a 2xx, moves on - and nobody has any reason to suspect anything is fragile. What is fragile is only the error path, which under calm traffic almost never fires.

It takes a spike to expose it. Several services send notifications at once, you hit the limit and get a 429, or Discord momentarily returns a 5xx. In the naive approach that event simply disappears: the code moves on, the exception at best lands in a log, and Discord is missing the entry that was supposed to be there. Worst of all, you lose it precisely at peak traffic, which is when notifications matter most.

That is why the proxy is not a performance optimization. It is a delivery guarantee under conditions where naive sending quietly fails.

1
Per-target queue

each webhook has its own queue, so one slow target does not block the rest.

2
Respect for 429

on a limit the proxy waits exactly as long as the Retry-After header says, not a guess.

3
5xx retries

a server error means exponential backoff and another attempt, not a lost event.

4
Hard stop on 4xx

a malformed payload will not become valid on the third try, so we do not waste attempts on it.

A queue that does not block the rest

The first decision is a separate queue for each target. If all events went through one shared queue, a single slow or blocked webhook would stall delivery to every other one - one clogged target and silence everywhere. Splitting the queues means a problem with one webhook stays with that one webhook.

The same structure gives resilience to spikes. When a wave comes in, events are not dropped - they land in their target's queue and wait for a window in which the limit lets go. From the outside this looks like a momentary delay instead of a loss, and that is an entirely different quality: a delayed notification still does its job, a lost one does none.

The heart: the delivery loop

All the logic comes down to one decision per response: success ends it, 429 says wait, 5xx says retry, other 4xx means stop without retrying. Telling those cases apart is the difference between a proxy that delivers and one that either drops events or grinds forever on a broken payload.

dispatch.ts · ts
async function dispatch(target: WebhookTarget, payload: Payload) {
  for (let attempt = 0; attempt < maxAttempts; attempt++) {
    const res = await post(target.url, payload);
    if (res.status < 300) return;
    if (res.status === 429) {
      await sleep(retryAfterMs(res));
      continue;
    }
    if (res.status >= 500) {
      await sleep(backoff(attempt));
      continue;
    }
    throw new PermanentError(res.status);
  }
}

The asymmetry: 429 versus 5xx

The key is the asymmetry between 429 and 5xx. On a 429 Discord itself says how long to wait - we read it from the header and wait exactly that, no less and no more. On a 5xx nobody says anything, so we add exponential backoff to avoid hammering a momentarily down server with a burst of immediate retries.

!
Warning

The most common mistake is retrying everything blindly. On a 429 a retry makes sense, but on a 400 or 404 it does not: the payload will not fix itself, and you only reach the limit faster and grind endlessly on something that will never go through. Retry what is transient, and only that. The rest should fail immediately, loudly and once.

The effect: a spike the other side never sees

For the services behind it the proxy is invisible - they send exactly as before and know nothing about limits or retries. All that complexity sits in one place instead of being smeared across every service separately. It also means a fix in limit handling goes in once, not into several codebases at once.

What the proxy does with a response (illustrative)

2xx delivered
92
429 wait and retry
5
5xx backoff and retry
2
4xx hard stop
1

The difference only shows at the next wave of traffic, and that is the whole point. Notifications do not vanish silently at the moment they are needed most - they wait in the queue and arrive once the limit lets go. One egress point turned silent event loss into at most a momentary delay, and that is the whole difference between monitoring you can trust and monitoring that goes quiet right when the fire starts.

More projects

More work from the same category - see how we tackle similar challenges.

Have a similar project?

Get in touch - a quote is free and comes back within an hour.