RAGKnowledge GraphLLMPython

One bad LLM response in 597 chunks: how FIRST_EXCEPTION failed our biggest documents

A malformed 200 from the model cancelled a whole batch of good extraction results, and bigger documents were far more likely to hit one. The fix and the pattern.

Daniel Voyce··11 min read

Our knowledge-graph extraction sends every chunk of a document to an LLM and collects the entities and relations that come back. In June 2026, a model response that arrived as a perfectly normal HTTP 200 with no content in it, roughly once in every 200 chunks, was enough to fail the graph extraction for an entire 597-chunk document, throwing away the results of every other chunk that was in flight alongside it.

One bad response could do that because of how the extraction loop waited on its chunks: asyncio.wait(..., return_when=FIRST_EXCEPTION), which treats a batch of independent units as one all-or-nothing job. Combine that with a retry that starts the batch again from its first chunk and you get a pipeline where the probability of finishing falls as the document gets longer. Our largest documents were the ones most likely to never get a graph at all.

How extraction fans out a document

Graph extraction in Certant runs on our fork of LightRAG. For each document, the chunks go through extract_entities in operate.py, which starts one asyncio task per chunk, limits how many run at once with a semaphore, and gathers the results. Above that, a wrapper in lightrag.py feeds a document's chunks through in checkpointed batches (50 chunks per batch by default), writing each finished batch to disk before moving on.

Here is the gather, as it still reads in our code and in upstream LightRAG:

tasks = []
for c in ordered_chunks:
    task = asyncio.create_task(_process_with_semaphore(c))
    tasks.append(task)

# Wait for tasks to complete or for the first exception to occur
# This allows us to cancel remaining tasks if any task fails
done, pending = await asyncio.wait(tasks, return_when=asyncio.FIRST_EXCEPTION)

...

if first_exception is not None:
    for pending_task in pending:
        pending_task.cancel()
    if pending:
        await asyncio.wait(pending)
    progress_prefix = f"C[{processed_chunks+1}/{total_chunks}]"
    prefixed_exception = create_prefixed_exception(first_exception, progress_prefix)
    raise prefixed_exception from first_exception

Read that with one failing chunk in mind. The first exception wakes the wait. Every task still pending is cancelled, including chunks that were halfway through a slow LLM call. Every task already in done that succeeded had its result appended to chunk_results, and that list is then abandoned, because the function raises instead of returning it. The batch wrapper catches nothing useful from that: it logs "Graph extraction batch N failed", doesn't save the batch, doesn't advance the checkpoint, and re-raises. The document's graph extraction is marked failed.

The comment in the code gives the reasoning: cancel the rest if anything fails. That's a sensible default when the tasks depend on each other and a partial answer is worthless. Chunk extraction doesn't work like that: each chunk is extracted independently, and a document with 596 chunks' worth of entities is far more useful than one with none.

The response that broke it

The exception came from a configuration file, which did this after every extraction call:

result = response.choices[0].message.content

Our inference providers occasionally return a 200 whose choices is None or an empty list. A content filter, an upstream hiccup or an error-shaped body can each produce it. We saw it at about 1 in 200 chunks. Subscripting None raises TypeError: 'NoneType' object is not subscriptable, which went up through _process_with_semaphore, picked up the chunk ID as a prefix, and landed in the FIRST_EXCEPTION wait above.

We already had a resilient path for bad extraction output, but it covered a different failure: a response that was truncated or missing its completion delimiter got up to three retries and then the pipeline carried on with the incomplete result. Nothing handled a response with no content at all, so that one crashed.

It surfaced while we were re-running 11 documents from a production knowledgebase whose graph extraction had failed. Eight of them re-merged cleanly. The other three were large documents that needed a full re-extraction, at 293, 336 and 597 chunks, and all three hard-failed on this exact TypeError. The 597-chunk document died in batch 5, at chunk 4 of 50.

Why the biggest documents were the ones that never finished

Once one chunk can sink the batch, the longer a document is, the more likely it is to fail, whatever it says. If every call has the same small chance of going wrong, the chance that at least one call in a run of N goes wrong is 1 − (1 − p)^N, and that climbs fast.

This table uses the 1-in-200 rate we observed for empty responses. It assumes each call fails independently, which I haven't measured, so treat it as an illustration of the shape rather than a forecast:

Chunks in one run Chance at least one call comes back empty
1 0.5%
10 4.9%
50 22.2%
100 39.4%
293 77.0%
336 81.4%
597 95.0%

At that rate, a 597-chunk document has about a 5% chance of getting through every call without a single empty response. Batch checkpoints soften this for long documents, since a failure only throws away the batch in flight, but each failure still fails the document and sends it back through the recovery path.

Empty responses weren't the only trigger. Anything that made one chunk's call raise had the same effect: a rate-limit error that outlasted the worker's retries, a timeout, a dropped connection, or a gateway routing setting that answered "No endpoints found" with a 404 that LiteLLM doesn't retry. During a bulk import on 19 June 2026 we had 119 documents fail graph extraction. Every one of them was transient: 30 explicit rate-limit errors, 3 timeouts, 3 connection errors, and about 83 that were the whole document aborted mid-extraction because one chunk's call failed. None were documents that couldn't be extracted.

The retry path made the size effect worse. Failed documents are picked up by a background sweep that re-dispatches them, bounded at 20 recovery attempts per document in a 24-hour window, so a broken document stops instead of looping forever. The checkpoint only records completed batches, and a document of 50 chunks or fewer is a single batch, so every retry started again at chunk 1. The extraction cache made the chunks that had already succeeded cheap on the second pass (they came back as cache hits rather than new calls), but the whole batch still had to run again and still had to get through every chunk cleanly.

We watched this happen on a set of large research documents, each 50 chunks. One of them extracted into the high 40s, hit a single failing call, failed, went back through the sweep, got close to the end again, and failed again, more than once. Most documents got a clean pass eventually and finished, but a large document with bad luck could use up all 20 attempts and end up permanently graph-failed.

It helped that the graph is a second stage for us. A document is chunked and embedded first, and becomes searchable as soon as that finishes, so these documents were answering questions through vector search the whole time. Only their contribution to the knowledge graph was stuck.

The fix we shipped: guard the response and degrade to empty

The fix went live on 22 June 2026 and has three parts, all in llm_provider.py.

First, a defensive accessor that never raises:

@staticmethod
def _safe_response_content(response) -> Optional[str]:
    try:
        choices = getattr(response, "choices", None)
        if not choices:
            return None
        message = getattr(choices[0], "message", None)
        if message is None:
            return None
        return getattr(message, "content", None)
    except Exception:
        return None

None from this function means "the provider gave us nothing usable", which the caller can act on, instead of a TypeError that the gather turns into a dead batch.

Second, the main non-streaming extraction call retries on an empty response, three times by default, and if every attempt comes back empty it degrades to an empty string:

result = None
for null_attempt in range(null_retries + 1):
    response = await client.chat.completions.create(**regular_params)
    result = self._safe_response_content(response)
    if result is not None:
        break
    # log a warning and retry
if result is None:
    # log an error: degrading to empty result so extraction can continue
    result = ""
return result

An empty string is safe here because the downstream code already copes with it. Plenty of chunks legitimately extract zero entities and zero relations, so an empty result flows through the parser and the merge without special handling. The document finishes minus that one chunk's entities.

The ten other places in the provider that read response.choices[0].message.content directly, on the text and vision fallback paths, now go through the same accessor with or "" on the end.

Degrading to empty is a deliberate trade-off: we can lose one chunk's entities without the document being marked as failed. We log it at error level with the knowledgebase ID so it can be found, and one chunk missing from a 597-chunk graph is a far better outcome than no graph at all. If the provider returns nothing four times running for the same chunk, I'd rather keep the rest of the document than keep retrying the whole thing.

The test file test_llm_null_response_guard.py has 12 tests covering the shapes a malformed response can take: choices=None, choices=[], a None response, an error body with no choices attribute, a choice whose message is None, a message whose content is None, a valid passthrough, and a parametrised set of junk inputs (an integer, a string, a bare object(), a dict with choices: None, an empty list) that must all come back as None without raising. Beyond the unit tests, we built the real provider in the worker container, patched the OpenAI client to return choices=None, and checked that the call made four attempts and returned "" without crashing. We ran the same check against the staging build and then the production build after deploying. The three large documents were then sent back for re-extraction.

Still to do: retry per chunk

The null guard removes one trigger. It doesn't change what happens when any other exception reaches that wait, and the FIRST_EXCEPTION gather is still in extract_entities. A rate-limit error that outlasts the retries, or a timeout, still cancels the batch and fails the document.

The structural fix is to stop treating the batch as the unit of success. Let every chunk task finish, keep the results that came back, and record which chunks failed. Save the good results and advance the checkpoint by chunk ID rather than by batch. On retry, extract only the failed chunks. The checkpoint already stores processed_chunk_ids, so the data model is there; what's missing is letting successful chunks commit when a sibling fails.

Here is the pattern in outline. It's a sketch and we haven't shipped it yet:

results = await asyncio.gather(
    *(_process_with_semaphore(c) for c in ordered_chunks),
    return_exceptions=True,
)

good, failed_ids = [], []
for chunk, res in zip(ordered_chunks, results):
    if isinstance(res, BaseException):
        failed_ids.append(chunk[0])
    else:
        good.append(res)

# persist `good` and mark those chunk IDs processed,
# then re-queue only `failed_ids` with their own retry budget

With that in place, a 597-chunk document that hits one bad call re-extracts that one chunk, and the recovery budget is spent on chunks that actually failed rather than on whole documents that were nearly done. For our goal of a few thousand documents ingesting unattended, this is the change that matters most.

Cancelling on the first exception does have one use under a rate limit: it stops the rest of the batch from piling more calls onto a provider that has just told you to back off. If you switch to letting everything finish, keep your backoff and your concurrency limit somewhere that still sees the failures, because otherwise every in-flight chunk hits the same wall one after another.

What I'd check in any long LLM batch pipeline

Find out what your concurrency primitive does when one unit fails, and decide whether that's what you want. asyncio.wait with FIRST_EXCEPTION hands you the first exception and leaves you to cancel the rest. asyncio.gather without return_exceptions=True raises the first exception and you never see the other results (the other tasks keep running, unobserved). A TaskGroup cancels everything still running. All three are right when the units only mean something together. For extraction, embedding, summarisation or classification over independent chunks they're wrong, and the table above shows why short test documents won't reveal it.

Check what's inside a 200 from a model API before you use it. Every field you subscript on a response object is a place a provider can hand you None. One accessor that returns None for anything malformed is cheap, and it lets your own code decide what happens next.

Make the checkpoint granularity match the retry granularity. If the smallest thing you can save is a batch of 50, the smallest thing you can retry is a batch of 50, and any document that fits in one batch restarts from scratch every time.

Plot failure rate against document size. If failures rise with chunk count and nothing else, some all-or-nothing step is sitting in the loop, and an aggregate failure rate hides it because small documents dominate the count.

If you want to see the output format side of the same extraction pipeline, I wrote about why asking for JSON halved our extraction recall.

Build a brain for your business.

Certant turns your documents, data and processes into agents, dashboards and assistants you can actually trust.