Two threads inserting into a full LRUCache hang forever on the GIL build (3.12, 3.14) when the values have a __del__. No gc.collect(), sleep or reentrant call is involved, so this is not the collector case from #84 that #86 fixed. With one thread, or without __del__, it finishes in about 0.1 s. On 3.14t it finishes, but hangs once a third thread runs gc.collect() every 10 ms.
import threading
import cachebox
cache = cachebox.LRUCache(100)
class Value:
def __del__(self):
self.dropped = True # any Python-level finalizer
def insert(base):
for i in range(200_000):
cache[base + i % 1000] = Value() # the cache is full, so this evicts an older Value
threads = [threading.Thread(target=insert, args=(b,)) for b in (0, 10_000)]
for t in threads:
t.start()
for t in threads:
t.join()
print("finished")
It never prints finished: killed by timeout 60 in 3/3 runs on each of 3.12.12 and 3.14.7. faulthandler shows one thread in Value.__del__ called from the insert line, and the other on the insert line.
cachebox 6.2.7 from PyPI (src/ unchanged on main at 9af1cec), CPython 3.12.12, 3.14.7 and 3.14.7t, Linux x86_64.
insert evicts while holding the cache's parking_lot mutex, and dropping the evicted value runs its __del__ under that lock. When __del__ gives up the GIL at the switch interval, the other thread takes the GIL and blocks in lock() without releasing it. On 3.14t the blocked thread stays attached, so a collection's stop-the-world pause waits for it forever.
|
pub fn insert( |
|
&self, |
|
py: pyo3::Python<'_>, |
|
handle: P::Handle, |
|
) -> pyo3::PyResult<Option<P::Handle>> { |
|
let mut lock = self.inner.lock(); |
|
self.insert_no_lock(&mut lock, py, handle) |
|
} |
|
fn evict(&mut self) -> pyo3::PyResult<()> { |
|
self.policy.evict(self.shared)?; |
|
Ok(()) |
|
} |
An untested idea: wait for the lock with PyO3's MutexExt::lock_py_attached (feature parking_lot), which detaches while blocked, or drop evicted values after the lock is released.
Found while testing with ftcheck.
Two threads inserting into a full
LRUCachehang forever on the GIL build (3.12, 3.14) when the values have a__del__. Nogc.collect(), sleep or reentrant call is involved, so this is not the collector case from #84 that #86 fixed. With one thread, or without__del__, it finishes in about 0.1 s. On 3.14t it finishes, but hangs once a third thread runsgc.collect()every 10 ms.It never prints
finished: killed bytimeout 60in 3/3 runs on each of 3.12.12 and 3.14.7. faulthandler shows one thread inValue.__del__called from the insert line, and the other on the insert line.cachebox 6.2.7 from PyPI (
src/unchanged on main at 9af1cec), CPython 3.12.12, 3.14.7 and 3.14.7t, Linux x86_64.insertevicts while holding the cache'sparking_lotmutex, and dropping the evicted value runs its__del__under that lock. When__del__gives up the GIL at the switch interval, the other thread takes the GIL and blocks inlock()without releasing it. On 3.14t the blocked thread stays attached, so a collection's stop-the-world pause waits for it forever.cachebox/src/policies/wrapped.rs
Lines 146 to 153 in 9af1cec
cachebox/src/policies/lrupolicy.rs
Lines 77 to 80 in 9af1cec
An untested idea: wait for the lock with PyO3's
MutexExt::lock_py_attached(featureparking_lot), which detaches while blocked, or drop evicted values after the lock is released.Found while testing with ftcheck.