PythonConcurrencyasyncioDebugging

Two locks that locked themselves: an fcntl.flock and a threading.Lock

Twice in seven weeks one thread asked for a lock it already held and froze our RAG service. Why, how we found it, and the reentrant fixes.

Daniel Voyce··11 min read

On 19 May 2026 the knowledgebases page in production came back empty and our RAG service reported itself unhealthy. One thread had taken an exclusive fcntl.flock on a metadata file and then, a few calls deeper, asked for the same lock again; the service runs on a single asyncio event loop, so every HTTP endpoint hung behind it, a health endpoint included, and the only thing that released the lock was restarting the process.

On 5 July I did it again with a different primitive. A config update held a plain threading.Lock and called get_config(), which took the same lock. Both times the bug had the same shape, a lock that isn't reentrant on a code path that can call back into itself. This piece covers both incidents and their fixes, and what I now look for when a function takes a lock and then calls something else.

How a thread waited on its own flock

Metadata writes in our RAG service go through a small set of helpers in common/file_utils.py: load_json_locked, write_json_locked and atomic_update_json_locked. The last one takes an exclusive lock on a sidecar .lock file, loads the JSON, hands it to a callback, and writes back whatever the callback returns. Each user has their own knowledgebase metadata file with its own lock file next to it.

MetadataManager.save_knowledgebase_metadata in a configuration file used that helper with a callback that builds the knowledgebase's config dict. Inside the callback, _build_config needs the user's saved defaults, and on one branch it fetched them by calling _load_metadata(). That calls load_json_locked on the same file, which opens the .lock file a second time and calls flock() on the new descriptor.

fcntl.flock locks are attached to the open file description, the kernel object you get back from open(), and not to the process or the thread. A second open() of the same file creates a second open file description, and the kernel treats it as a separate contender. So the same thread, in the same call stack, sits waiting for a lock that it holds itself, and it will wait forever, because the code that would release the first lock is further up the stack that is now blocked.

Stripped down to the shape (this is a sketch, not our code):

import fcntl

def load_json_locked(path):
    with open(path + ".lock", "w") as lk:           # second open file description
        fcntl.flock(lk.fileno(), fcntl.LOCK_SH)     # blocks: the EX lock below is held
        ...

def atomic_update_json_locked(path, update_func):
    with open(path + ".lock", "w") as lk:
        fcntl.flock(lk.fileno(), fcntl.LOCK_EX)     # first lock, held
        data = ...
        update_func(data)                           # callback calls load_json_locked(path)
        ...

Our RAG service handles HTTP on a single asyncio event loop, and this code was synchronous. The blocked flock() call was therefore blocking the loop itself, and every endpoint on the service hung with it: the knowledgebase listing returned nothing and a health endpoint timed out.

Why the bug first fired on 19 May

The code already had a guard for this exact case. The save callback has the loaded file contents in hand, so it pulled the user's _defaults out of them and passed them into _build_config as _preloaded_defaults. _build_config was only meant to go back to disk when nothing had been passed.

The guard tested truthiness. A brand-new per-user metadata file has no _defaults key yet, so the callback passed {}. An empty dict is falsy, the guard read it as "nothing was provided", and _build_config re-read the file under the lock its own caller was holding.

So the deadlock needed the first knowledgebase save into a fresh per-user metadata file. That path had never run in production. In the previous release we had fixed a Save action that had been a silent no-op, and once Save did something, the first write to a fresh per-user file walked straight into it.

Finding it took a look from outside the process, because from inside nothing was moving. Every thread's /proc/<pid>/task/*/wchan read locks_lock_inode_wait, which is where a task parks while it waits on a file lock, and the process had two file descriptors open on the same .lock file. Those two readings together were enough to identify the bug without attaching a debugger.

Recovery was a container restart. flock locks are advisory and the kernel frees them when the holding process dies; nothing inside the process was ever going to release this one.

The sentinel and the reentrant flock layer

The narrow fix is the guard. _build_config now distinguishes "not provided" from "provided and empty" with a sentinel instead of a truthiness test:

_DEFAULTS_UNSET = object()

def _build_config(self, kb_id, name, description, parser,
                  existing_created_at=None,
                  _preloaded_defaults=_DEFAULTS_UNSET, **kwargs):
    if _preloaded_defaults is _DEFAULTS_UNSET:
        # Not called under the metadata file lock: safe to load defaults.
        all_data = self._load_metadata()
        ...
    else:
        # Explicitly provided (possibly {}). We are inside the held file
        # lock: never re-read the same file.
        user_defaults = _preloaded_defaults or {}

That closes the trigger we hit. It does nothing for the next callback someone writes that happens to read the file it is being called under, so the bigger change was making the locked helpers reentrant for the same thread.

The layer keeps a process-wide registry of the flocks this process holds, keyed by (os.path.realpath(lock_file), threading.get_ident()), with a depth counter per entry. When a thread asks for a lock it already holds on the same file, the helper bumps the depth and reuses the descriptor it already has instead of opening a second one. The real LOCK_UN and close() happen only when the depth drops back to zero. The core of it, from common/file_utils.py:

key = (os.path.realpath(lock_file), threading.get_ident())

with _REGISTRY_LOCK:
    existing = _HELD.get(key)
    if existing is not None:
        existing.depth += 1          # same thread, same file: reuse the held fd
        reentrant = True
    else:
        reentrant = False

if not reentrant:
    fileobj = open(lock_file, "w")
    fcntl.flock(fileobj.fileno(), flock_mode)   # blocking call, outside the registry mutex
    with _REGISTRY_LOCK:
        _HELD[key] = _Held(fileobj, mode)

The registry mutex is never held across the blocking flock() call, because doing that would serialise every lock on every file behind whichever one was slowest. Only same-thread recursion is collapsed: a different thread still takes its own real flock and gets proper mutual exclusion, and a different process still serialises through the kernel. And there is one case the layer can't do safely. Upgrading a shared lock to exclusive while it is held can't be done on the same open file description, so if that ever happens the layer reuses the shared lock and logs a warning that names the file and says there may be a lost update. Our code paths don't do it; the warning is there so a test finds it if they ever start.

The same helpers exist twice in the repository, once in common/file_utils.py and once in the LightRAG code we carry (a configuration file), and both got the reentrant layer.

The third change was about the event loop. Some storage classes took a synchronous flock while holding an async storage_lock, so a legitimate wait on another process (cross-process contention, as opposed to a deadlock) would still stall the loop. Those call sites in json_kv_impl.py, json_doc_status_impl.py and networkx_impl.py now run the blocking part in asyncio.to_thread. We deliberately left one synchronous: the graph index_done_callback read-modify-write in networkx_impl.py. It runs in its own graph-merge worker process at concurrency 1, isn't on the request loop, can't self-deadlock, and a mistake there would corrupt the graph.

How we checked it

Each test that could deadlock runs under a watchdog thread, so a regression shows up as a failure in seconds instead of a hung test runner. The numbers below come from the project's record of the fix and the staging deploy.

Check Result
Unit suite maintained_tests/test_reentrant_flock.py 5 of 5 pass; the exact nested-load case went from an infinite hang to 0.00 s
Same suite, cross-thread and cross-process cases mutual exclusion between threads and serialisation between processes both preserved
Real MetadataManager inside the service container 3 of 3: first save into a fresh per-user file in 0.01 s, 10-way parallel saves, no lost updates
Parallel ingest of real PDFs, file backend 28 chunks on disk
Parallel ingest of real PDFs, PostgreSQL backend 28 chunks and 192 entities
a health endpoint polled during about 40 minutes of cumulative parallel ingest on both backends 2,504 of 2,504 and 3,551 of 3,551 answered, maximum 38 to 114 ms; the loop never froze
Staging, after deploying the fix fresh-user save 0.00 s, 8-way parallel pass, a health endpoint 200

The fix was tagged and running on staging on 20 May 2026, the day after the incident.

A threading.Lock in the config manager, 5 July

We had been adding a model alias layer in LiteLLM, because changing the model behind anything meant editing four places: a system config that was cached in memory, per-knowledgebase configs, code defaults, and the gateway config. On 5 July I tried to change the production answer model through the RAG service's own internal config endpoint.

The request hung. Within seconds every get_config() call in the service blocked behind it, the health check got no response at all, and knowledgebase listing, counting and querying all froze. The parts of the product served by a separate agent API carried on as normal.

A release had deployed shortly before, so that was the first suspect. The release had been healthy, though, and had served dozens of good queries after it went out.

The config manager in a configuration file guarded its state with self.config_lock = Lock(). The update methods looked like this:

def update_fast_llm_config(self, fast_llm_config):
    ...
    with self.config_lock:
        config = self.get_config()          # takes config_lock again
        config["fast_llm_config"] = config_to_save
        if self._save_config(config):
            self._config_cache = config
            return True
        return False

def get_config(self):
    with self.config_lock:
        if self._config_cache is None:
            self._config_cache = self._load_config()
        return self._config_cache.copy()

update_llm_config and update_embedding_config had the same pattern. A threading.Lock doesn't track which thread owns it, so a second acquire() from the owning thread blocks like any other caller. The update thread hung while holding the lock, and every request that read config queued behind it: queries, knowledgebase listing and counting all do. Restarting the container cleared the held lock and everything came back with its data intact.

The fix is one word: self.config_lock = RLock(). An RLock records its owning thread and a count, so a re-acquire from the owner increments the count and returns straight away, and the lock is released when the count gets back to zero. The singleton _lock in the same class, which only guards instance creation, stays a plain Lock because nothing re-enters it. One method, update_workflow_builder_models, had already been written to call get_config() outside the lock. We kept it that way even though it would now be safe inside, because holding the write lock for the shortest possible window is still better.

Four regression tests in a configuration file cover it. The first asks the live config manager's lock to be acquired twice, non-blocking, from the same thread, which fails immediately if anyone swaps it back to a plain Lock. Two run an update followed by a read under a 10-second watchdog. The last runs a reader and an updater thread against each other for 20 iterations each and asserts that neither wedges and that all 20 reads complete. On staging, the same config POST that had hung for two minutes returned in 0.01 s, and the service stayed responsive throughout.

One fix was a single word and the other needed a registry. Python ships a reentrant thread lock. There's no reentrant flock, because the kernel treats the open file description as the owner, so a thread that opens the file again counts as a separate contender. For the file lock we had to build the bookkeeping RLock does for you: an owner (file plus thread) and a count.

This is a different lock problem from the one in 155 pages of OCR took 37 minutes. The bottleneck was one threading.Lock.. That lock worked as designed and serialised the work behind it; the two locks here were never released.

The operational lesson from July is separate from the code one. A config change should not be able to wedge the whole service, and the bug that let it was real, but I also shouldn't have been trying out a config endpoint on production. Config changes go through staging first now.

What I look for now

Any function that takes a lock and then calls something else gets a second look, and callbacks get the hardest look of all. atomic_update_json_locked couldn't see what its callback did; the callback's author couldn't see from inside _build_config that a lock was held two frames up. Each piece looked fine read on its own, and the bug stayed hidden until a new path joined them up.

If the lock is a threading.Lock and there is any chance the protected section calls back into code that takes it, I use an RLock. The cost is a little bookkeeping. For file locks there's no standard library answer, so a reentrant wrapper keyed by file and thread is worth having before you need it, and it has to leave cross-thread and cross-process exclusion alone or it becomes a different bug.

"Was this argument provided?" gets a sentinel, never a truthiness check. An empty dict is a legitimate value, and it looks exactly like "missing" to if not x.

On an asyncio service, a synchronous blocking lock can freeze every request on the loop along with the one that took it. Where a legitimate wait is possible, the blocking call goes through asyncio.to_thread.

Deadlock tests run under a watchdog with a short timeout, so a regression fails in seconds with the test's name instead of hanging the whole run.

When the only recovery is a process restart, that points at a lock nobody will ever release. Read wchan and the open descriptors before you restart, while the evidence is still there.

Build a brain for your business.

Certant turns your documents, data and processes into agents, dashboards and assistants you can actually trust.