Four data-loss bugs that came from one dict merge
Four vector-store bugs lost up to 71% of chunks and entities without raising an error. All four trace back to one unguarded write-behind merge.
Between 5 and 7 January 2026 we traced four data-loss bugs in the vector storage layer of our ingestion pipeline. They lost 38 to 65% of chunks on multi-document uploads, overwrote 2,691 entity IDs within a single document, mixed five documents' chunks into one merge, and left up to 71% of a document's entities in memory only, so they vanished on restart.
None of them raised an exception, and every document finished as processed. All four came down to one line of Python in the vector store's write path, and how it was called:
merged_data = {**disk_data, **in_memory_data}
This piece is about that line, and the checks a write-behind store needs so that you find out when it has thrown records away.
How the write-behind store merges memory onto disk
Our vector storage is the NanoVectorDB backend we inherited with our LightRAG fork, and it has three namespaces: chunks, entities and relationships. It's a write-behind design. upsert() puts records into an in-memory client, and later index_done_callback() persists them. The callback loads what is on disk, builds two dicts keyed by record ID, merges memory on top of disk, rebuilds the vector matrix and writes the file back:
disk_data = {item['__id__']: item for item in disk_storage.get('data', [])}
in_memory_data = {item['__id__']: item for item in in_memory_storage.get('data', [])}
# Dict merge: overlapping keys are overwritten, not accumulated
merged_data = {**disk_data, **in_memory_data}
On an ID collision a Python dict merge keeps the right-hand value and discards the left. The merged dict comes out shorter than the two inputs added together, and no error is raised. That's fine when a collision means "the same record, updated". It's a disaster when a collision means "a different record that happens to share a key", or "a record that was already on disk and has been dragged back into memory by mistake".
The design only holds if two things are true. Memory must contain exactly what this write is adding, and the callback must run. Every bug below breaks one of those assumptions. Three of them put the wrong things in memory; the fourth never flushed memory at all.
Bug 1: a worker's memory was never emptied between documents
Celery reuses worker processes. When one worker handled several documents for the same knowledgebase in a row, the in-memory client was never cleared after a successful write. Each new document's chunks landed on top of every previous document's chunks, still sitting in memory.
The first sign was a log line whose arithmetic did not work:
Process 20 merged: 420 disk + 589 memory = 589 total (0 preserved from concurrent writes)
Rebuilt matrix with 319 vectors (should be 1009)
420 plus 589 should be 1,009. All 420 disk records had the same IDs as 420 of the memory records, so the merge overwrote them rather than adding to them, and the matrix rebuild then kept only the entries that still carried a vector.
Walk it through with three documents processed in sequence by one worker: A with 254 chunks, B with 252, C with 426. After A, memory still holds A's 254. B arrives and memory holds 506. By C, memory holds all 932 and disk holds A and B, so almost everything collides. The per-document counts before the fix, from the fix write-up (the first document is Aruma's enterprise agreement, the other two are enterprise agreements from the same test set):
| Document | Expected chunks | Persisted before fix | Loss |
|---|---|---|---|
| First (Aruma) | 254 | 254 | 0% |
| Second | 252 | 156 | 38% |
| Third | 426 | 150 | 65% |
The first document in any run was always fine, so a test with one document would never show it. The fix, on 5 January 2026, was to empty the in-memory data list and reset the matrix (keeping its dtype and embedding dimension) at the end of every successful write, and to log how many entries were freed. The expected result after it was 254, 252 and 426 chunks on disk, with a merge line reading 506 disk + 426 memory = 932 total.
Bug 2: the same entity ID written twice within one large document
Clearing between documents didn't help a large document, because a large document is processed in batches inside a single worker task. Batch one extracts a run of entities and calls index_done_callback(). Batch two extracts the next run and calls it again.
Entity IDs are deterministic: ent-doc-{doc_id}-{entity_hash}. An entity mentioned early and late in a long document (an "Employee" in chunks 1 to 500 and again in 501 to 1,000) gets the same ID in both batches. So the second batch's memory contains IDs that the first batch has already written to disk, and the merge overwrote them.
The logs for one large enterprise agreement, from 7 January 2026:
34 calls to index_done_callback for entities (one document!)
4 calls with 2691 vectors in memory
Disk progression: 8435 → 11126 entities
CRITICAL: Found 2691 overlapping IDs! All from SAME document
MERGE MATH ERROR: Expected 13817 (11126 disk + 2691 memory) but got 11126
Our validation script reported that document at 50% entity loss, and the investigation classed it as real loss.
The same stale memory produced a second, stranger symptom. The verification step after each write took the document's file path from the first in-memory record that had one, then counted that document's vectors on disk. If memory still held a record from the previous document, verification counted the wrong document. One upload merged correctly (14229 disk + 2955 memory = 17184 total), then failed verification with Only 1782/2955 vectors found because it had gone looking for another document's file path and found that document's 1,782 entities. Across the seven large documents we checked, one showed 100% loss that was entirely a false positive, another showed 68% that was probably false, and one at 8% was a mix of real and false. So stale memory was losing data and also making our loss detector unreliable, which is a bad combination when you are trying to work out how much you lost.
Two changes landed for this. The merge now computes the overlap explicitly, keeps the disk version of any duplicate (from the earlier batch), logs the count, and merges only the new records. And upsert() records the IDs in the current batch, so verification takes its file path from those IDs rather than from whatever happens to be first in memory.
Bug 3: five documents' chunks in one worker's memory
Uploading several small numbered test documents at once produced this:
CRITICAL: Found 25 overlapping IDs! Sample overlapping IDs from MULTIPLE documents:
- chunk-doc-8ec7746fdd99fa73509b365c988f7c6c
- chunk-doc-bd8e64db331918e2d3eff0434fb3d720
...
MERGE MATH ERROR: Expected 50 (25 disk + 25 memory) but got 25
Verification checking WRONG file
One worker's in-memory store held chunks from five different documents at once.
Our first suspect was a freshness flag. storage_updated exists so that a process reloads the client from disk when another process has written, which keeps query results current. _get_client() checks it, and upsert() calls _get_client(). The theory was that document A's callback set the flag, document B's upsert then reloaded every chunk on disk into memory, and B's chunks were added on top of A's. The story was plausible, since a reload meant for querying has no business running inside a write, but the logs didn't support it. We tried two fixes on that theory. Clearing storage after callback failures came too late, because the contamination had already happened. Clearing storage after a reload in _get_client() never ran, because no reload log line ever appeared.
The logs that settled it showed process 33 with 12 vectors in memory, 6 of them from a different document, before any reload logic could have run. The contamination came from __post_init__. When a new worker process creates its storage instance, it loads the whole vector file from disk, which contains every previously processed document. That load became the starting in-memory state, and the next upsert added to it. The fix, on 6 January 2026, clears data and the matrix immediately after the initial load and logs it:
Process 33 cleared 6 entries from initial chunks load to prevent cross-document contamination
Process 33 has 6 vectors in memory for chunks
Disk has 6 vectors for chunks
Process 33 merged: 6 disk + 6 memory = 12 total
VERIFICATION PASSED: All 6 vectors persisted correctly
Disk data isn't lost by doing that. The callback reloads disk at write time anyway and merges against it, so memory only ever needs to hold what the current document adds.
Bug 4: the entity and relationship stores were never flushed
This one was found first, on 5 January, and it's the simplest of the four, because the merge never ran at all.
Entities were upserted into their vector store at four call sites in operate.py. Chunks were persisted by chunks_vdb.index_done_callback(). The key-value stores for entities and relations had their callbacks. The entity and relationship vector stores, built the same way, had none:
$ grep -n "await.*entities_vdb.*index_done\|await.*relationships_vdb.*index_done" rag_service/lightrag/*.py
# no output
So that data sat in memory until something else happened to flush it, or the process died. The evidence was a count comparison between two stores that should agree. For one document, the key-value store had all 2,951 extracted entities (its callback ran). The entity vector store had 845 of them, a 71% loss.
Before finding this we'd spent months on file locking for these stores: fcntl.flock(), atomic writes through a temp file and rename, asyncio locks, merge-on-write and a serialised write queue. A knowledgebase built before the locking work showed 58% entity loss and 48% relationship loss. One built after it still showed 28.5% and 13.2%. The locking was working; the loss that remained came from a write that was never issued.
The loss varied between knowledgebases because whatever did reach disk got there by accident: a shutdown path, a later operation that triggered a write while the process was still alive, or leftover data from an earlier run. Loss that depends on timing like that looks like a concurrency problem, and concurrency was where we had been looking.
The fix was two awaited calls after the graph merge. In the code today, _insert_done() flushes every storage from one list, entities_vdb and relationships_vdb included, so the vector stores go to disk alongside everything else. The recovery plan needed no re-extraction: the extraction results were still in the LLM response cache and the entity key-value store, so only the merge phase had to be re-run for each affected document.
What changed, and the rule I'd apply to any write-behind store
The storage layer now clears memory after every successful write and straight after the initial load from disk, detects ID overlaps explicitly and keeps the disk version, takes the verification target from the current batch's IDs, and flushes every storage from one list. Chunk IDs are also namespaced by document (chunk-{doc_id}-{hash}), so two documents can't collide on a chunk ID at all.
The part I'd copy into any other system is the arithmetic check. The merge now computes what the result should be and compares:
expected_count_deduplicated = disk_count + in_memory_count - len(overlapping_ids)
if merged_count != expected_count_deduplicated:
logger.error(
f"MERGE MATH ERROR: Expected {expected_count_deduplicated} "
f"({disk_count} disk + {in_memory_count} memory - {len(overlapping_ids)} duplicates) "
f"but got {merged_count}. This indicates unexpected data loss!"
)
Three of the four bugs showed up in merge arithmetic in the logs. 420 disk + 589 memory = 589 total is a log line that says, in plain numbers, that 420 records just disappeared, and it is easy to miss when the document still finishes as processed. That is the same blind spot as an ingestion success check that only asks whether there are any chunks, which I wrote about in Asking for JSON halved our extraction recall.
The general rule I'd take from this is that a write-behind store should be able to account for every record it accepted. For the merge, that means the output count equals disk plus memory minus the duplicates you expected, and anything else is an error. Our check logs at error level rather than raising. Given what these four bugs cost, I'd want it to fail the write. A collision you didn't plan for should stop the pipeline, because {**a, **b} keeps the right-hand value without logging anything. For the flush, it means an independent count to compare against. The entity key-value store is what exposed the missing callback, and without count-based validation across the two stores we would not have found it.
If you run anything with an in-memory layer over a file or a table, grep for dict merges and update() calls on the persistence path and check what each one does when a key is already there.


