LLM OpsLiteLLMRate LimitsRetries

Three instant retries turned one rate limit into an overnight 429 storm

An overnight ingest logged 307.2K rate-limit errors while the provider's concurrency sat flat. The gateway's instant retries were the cause.

Daniel Voyce··9 min read

During an unattended overnight bulk ingest on 19 June 2026, the gateway's logs recorded 307.2K HTTP 429 responses, peaking at 51.56K in the 07:00 hour. The upstream model provider's own dashboard showed our concurrent requests sitting flat at about 186 all night, under the 200 we're allowed, which is not what a rate-limited account is supposed to look like.

The cause was one line in the LiteLLM config: num_retries: 3. LiteLLM's router fires those retries immediately, with no backoff, so every 429 became four requests in the same instant against an endpoint that had just said it was full. The fix was changing that 3 to a 1. Documents finishing graph extraction went from 5 to 17 per five minutes to 86, and the 429 count after the change was zero.

That 3 was the wrong value for how we were set up at the time, and working out why shows where retry logic should live when a gateway sits between your workers and a rate-limited model API.

The setup

Our ingest pipeline extracts a knowledge graph from every chunk of every document with gpt-oss-120b, served by an upstream model provider. Those extraction calls are token-heavy (around 108 s each) and there are a lot of them. The goal of the run was a bulk import that completes unattended, the sort of job we want to be able to start at 2,000 to 5,000 documents and walk away from.

Between the extraction workers and the provider sits LiteLLM as a gateway. It does two jobs that matter here. It caps concurrency: the provider allows 200 concurrent requests per model, and the gateway holds its own in-flight count at 180 using an asyncio semaphore per deployment. When that semaphore is full, callers queue instead of receiving a 429, which gives the workers clean backpressure:

router_settings:
  default_max_parallel_requests: 180

It also retries. The relevant line in litellm_settings was:

num_retries: 3

Behind both of those, the worker has its own retry layer: a tenacity decorator on the completion call that retries rate limits, timeouts, connection errors and 5xx responses with exponential backoff.

In the same round of work we'd also turned the LLM extraction cache back on (it had been silently off, which is its own story), so repeated chunks and resumed documents became cache hits instead of fresh calls. That made the workers faster, and against a hard ceiling the extra speed came back as more 429s.

Two wrong theories

The first theory was that the provider was slow. Throughput had dropped and errors were up, so it was the obvious suspect. The provider's dashboard ruled it out: a flat ~186 concurrent requests the whole night, so nothing had changed on the supply side.

The second theory was that the gateway had come up with the wrong number of worker processes. The semaphore behind default_max_parallel_requests is per process, not global. With one gateway process, 180 is the effective cap; with two, each process gets its own 180 and the real cap becomes 360, well past the provider's 200. That would have explained the 429s, but the gateway was running a single process with no restarts all night, so the effective cap was a flat 180 for the whole run.

Both theories were about concurrency, and concurrency was fine. Our own in-flight count from the TCP side (established connections) was 140 to 180, consistent with the provider's number. The provider's gauge sat at 186, below the 200 limit, because a request the provider rejects with a 429 never registers as concurrent at all. The limit we were hitting was the request rate, and something was pushing the rate up without pushing concurrency up.

The gateway's logs broke the night's errors down like this:

Status Count Notes
429 307.2K cumulative peak 51.56K in the 07:00 hour
499 3.16K
500 13

Every 429 was on one model group, openai/gpt-oss-120b. Every /embeddings call over the same period came back a clean 200. Embeddings use a different model with its own limit and had plenty of headroom, which put the problem squarely on the chat model we were hammering.

Instant retries against a full bucket

A 429 means "not now", and the polite response is to wait. LiteLLM's router, with num_retries: 3, fires the retry straight away instead, and if that gets a 429 it fires again, up to three times. So one rate-limited request turns into four requests in a burst, all aimed at the endpoint that has just told you it's at its limit.

That burst raises the request rate, which produces more 429s, which produce more instant retries, and nothing in the loop spaces anything out. Once the pipeline was running at the ceiling it stayed pinned there, spending requests on retries that had almost no chance of succeeding.

The two retry layers multiply. The worker's decorator allows 5 attempts. Each of those attempts goes through the gateway, which with num_retries: 3 makes up to 4 upstream requests. That's up to 20 upstream requests for one logical extraction call, and 15 of them (3 per worker attempt) are fired with no gap at all. With num_retries: 1 the ceiling is 10, and only 5 of those are instant.

The documents that failed during the storm were all transient failures. Of 119 failed documents, 30 were explicit 429s, 3 were timeouts, 3 were connection errors, and about 83 were documents aborted mid-extraction when a single chunk's call hit a 429 and took the whole document down with it (that pattern gets its own write-up in One bad LLM response in 597 chunks). None were permanent rejects. A sweep that re-dispatches failed documents on a bounded budget picked them up once the storm stopped, and the warm extraction cache made the re-runs cheap.

Changing num_retries from 3 to 1

We changed the gateway's retry count from 3 to 1:

litellm_settings:
  num_retries: 1

Applying it took a gateway restart and didn't disrupt the workers. After the change:

Measure Before After
Documents completing graph extraction per 5 minutes 5 to 17 86
429 responses 307.2K over the night 0

We left the worker concurrency alone. Turning down the number of concurrent extraction tasks feels like the direct response to a rate limit, but I think of worker concurrency as a lever and the gateway as the governor: the gateway decides how many requests reach the provider and what happens when one bounces. The governor was the part that was misconfigured, so that's where the fix went.

We kept num_retries at 1 instead of 0 on purpose. A single immediate retry still covers one-shot transients like a dropped socket, where trying again straight away is the right call, without re-bursting a rate limit.

What good backoff for a rate limit looks like

With one upstream provider, the gateway's instant retries were pure fuel. The retry that should absorb a 429 is the one that spaces requests out, and we already had it in the worker:

@retry(
    stop=stop_after_attempt(5),  # 5 attempts for transient errors
    wait=wait_exponential(multiplier=1, min=4, max=30),  # Exponential backoff
    retry=(
        retry_if_exception_type(RateLimitError)           # 429 - Rate limit
        | retry_if_exception_type(APIConnectionError)     # Network/connection issues
        | retry_if_exception_type(APITimeoutError)        # Request timeout
        | retry_if_exception_type(InternalServerError)    # 500 - Server error
        | retry_if_exception_type(InvalidResponseError)   # Invalid response format
        | retry_if_exception(is_transient_bad_request)    # 400 - All BadRequestErrors
        | retry_if_exception(is_server_error)             # Any 5xx server error
    ),
)
async def openai_complete_if_cache(...):

In tenacity, wait_exponential computes multiplier * 2 ** (attempt - 1) and clamps it between min and max. With min=4 the floor dominates: the waits between our five attempts are 4, 4, 4 and 8 seconds, about 20 seconds of spacing in total, and the 30 second cap is never reached. The embedding function uses the same shape with max=60. That schedule has no random jitter in it; we haven't measured whether adding some would change anything for this workload, so I won't claim it would.

Out of this incident, and the config changes around it, came a short set of rules we now follow.

One layer owns the spacing. The worker's retry is the one with backoff, so the worker is the layer that should handle a 429. A second layer that retries instantly underneath it multiplies the request count at exactly the moment the provider is asking for fewer. If you have a gateway with its own retry setting, find out whether it backs off before you trust it with rate limits.

Timeouts are ordered so the owning layer fails first. The worker's HTTP read timeout is 180 s, and the gateway's request_timeout is 240 s, one notch higher. The worker times out first and decides what to do next through its own backoff. The gateway's timeout is a backstop for the case where the worker's disconnect doesn't propagate. This ordering came from a separate incident the day before, where a request got lost inside the gateway and never returned. With the old 600 s default at both ends, that held a worker for up to 600 s per attempt, around 50 minutes across five attempts, and wedged graph extraction.

Concurrency is capped where the queue can wait. The gateway semaphore turns "too many in flight" into callers waiting their turn instead of the provider returning 429s. That's the right place for concurrency control, because waiting in a local queue never costs an upstream request. It doesn't protect you from a rate limit made worse by instant retries, because rejected requests don't count towards concurrency. The semaphore and the retry policy have to be set together.

The right retry count depends on topology. With a single provider behind a model name, instant gateway retries all land on the same saturated pool, so they amplify. With two or more independent pools behind the same model name, an instant retry after a 429 can land on a different pool that has room, and then it's useful failover.

Our config has since moved to that second shape. We added independent pools for gpt-oss-120b under the same model_name, and num_retries is back at 3. The comment above it in the config now reads:

  # 3 retries for cross-provider FAILOVER. Safe here because router_settings.
  # disable_cooldowns (below) keeps every deployment in rotation, so a retry after a
  # 429 lands on a DIFFERENT provider/pool (useful spillover) rather than re-bursting
  # the SAME saturated pool (the no-backoff amplifier that bit us with one provider).
  num_retries: 3

That only holds because of disable_cooldowns: true. LiteLLM's default behaviour pulls a deployment out of rotation after a burst of failures, and with every pool cooled at once the model group can go empty. With cooldowns disabled, a 429 just moves the retry to the next pool and a bad pool wastes one attempt. The same num_retries: 3 that caused the storm with one provider is the correct value with several, and the worker's exponential backoff is still the outer safety net either way.

The question I now ask about any gateway's retry setting is "when this retries, where does the request go, and how long does it wait first?". If the answers are "the same place" and "no time at all", the right number is close to zero.

Build a brain for your business.

Certant turns your documents, data and processes into agents, dashboards and assistants you can actually trust.