Is this a duplicate?
Type of Bug
Something else
Component
cuda.core
Describe the bug
Consider two threads entering deviceptr_import_ipc in _ipc.pyx, thread A and B. This call holds the gil.
Thread A is first and encounters this section of code in resource_handles.cpp:
std::lock_guard<std::mutex> lock(ipc_import_mutex);
if (auto h = ipc_ptr_cache.lookup(key)) {
return h;
}
GILReleaseGuard gil;
Then:
- Thread A takes the mutex.
- Thread A Releases the gil.
- Thread B gets the gil and gets to the mutex lock. It is not blocked while holding the gil.
- Thread A reaches the end of scope and does
~gil before releasing the mutex due to the ordering here.
So now hread B has the gil and wants the mutex. Thread A has the mutex and is trying to reacquire the gil. It's a deadlock.
Suggested fix
Just reorder GILReleaseGuard above the lock_guard in deviceptr_import_ipc, so the GIL is released before the mutex is taken and re-acquired only after it has been dropped.
Found this while reviewing #2759, which changes the cleanup function from fprintf(stderr, ...) to PyErr_WarnExrequireing the gil. Auditing where reports can be emitted from surfaced deviceptr_import_ipc as the one pw_* call site inside a held C++ lock.
How to Reproduce
Reproducer
Four threads each import a distinct descriptor, so all of them take the slow path. Every import succeeds — nothing fails and nothing is reported — which is what shows the hang is purely the lock ordering rather than anything on an error path.
import faulthandler
import multiprocessing as mp
import threading
THREADS = 4
BUDGET_SEC = 30
def child_main(queue, results):
# If we wedge, dump every thread's stack and abort, so the hang is visible.
faulthandler.dump_traceback_later(BUDGET_SEC, exit=True)
from cuda.core import Buffer, Device
device = Device()
device.set_current()
mr = queue.get()
descriptors = queue.get()
done = [0] * THREADS
barrier = threading.Barrier(THREADS)
def importer(index):
# A current context, so the import succeeds and nothing is reported.
Device().set_current()
barrier.wait()
try:
buffer = Buffer.from_ipc_descriptor(mr, descriptors[index], stream=device.default_stream)
done[index] = 1
buffer.close()
except Exception as exc:
done[index] = f"raised {type(exc).__name__}"
threads = [threading.Thread(target=importer, args=(i,), daemon=True) for i in range(THREADS)]
for thread in threads:
thread.start()
for thread in threads:
thread.join(timeout=BUDGET_SEC)
results.put({"alive": [t.is_alive() for t in threads], "done": done})
if __name__ == "__main__":
mp.set_start_method("spawn", force=True)
from cuda.core import Device, DeviceMemoryResource, DeviceMemoryResourceOptions
device = Device()
device.set_current()
if not device.properties.memory_pools_supported:
raise SystemExit("device does not support mempools")
mr = DeviceMemoryResource(device, DeviceMemoryResourceOptions(max_size=2 << 20, ipc_enabled=True))
stream = device.default_stream
buffers = [mr.allocate(64, stream=stream) for _ in range(THREADS)]
stream.sync()
queue = mp.Queue()
results = mp.Queue()
proc = mp.Process(target=child_main, args=(queue, results))
proc.start()
queue.put(mr)
queue.put([b.ipc_descriptor for b in buffers])
try:
outcome = results.get(timeout=BUDGET_SEC + 30)
print("child outcome:", outcome)
if any(outcome["alive"]):
print("DEADLOCK: importer threads never returned")
else:
print("no deadlock: every import completed")
except Exception:
print("DEADLOCK: importing process produced no result, and no report was ever emitted")
proc.join(timeout=10)
print("child exitcode:", proc.exitcode)
if proc.is_alive():
proc.kill()
Expected: four successful imports, exit 0.
Actual: the importing process never completes and is killed by its own watchdog.
Timeout (0:00:30)!
Thread 0x000076f4bffff6c0 [Thread-1 (impor] (most recent call first):
File "repro.py", line 72 in importer # <- inside Buffer.from_ipc_descriptor
...
Thread 0x000076f4bf7fe6c0 [Thread-2 (impor] (most recent call first):
File ".../threading.py", line 740 in wait
File "repro.py", line 70 in importer # <- woken, cannot re-acquire the GIL
...
DEADLOCK: importing process produced no result, and no report was ever emitted
child exitcode: 1
Expected behavior
No deadlock.
Operating System
No response
nvidia-smi output
No response
Is this a duplicate?
Type of Bug
Something else
Component
cuda.core
Describe the bug
Consider two threads entering deviceptr_import_ipc in _ipc.pyx, thread A and B. This call holds the gil.
Thread A is first and encounters this section of code in resource_handles.cpp:
Then:
~gilbefore releasing the mutex due to the ordering here.So now hread B has the gil and wants the mutex. Thread A has the mutex and is trying to reacquire the gil. It's a deadlock.
Suggested fix
Just reorder
GILReleaseGuardabove thelock_guardindeviceptr_import_ipc, so the GIL is released before the mutex is taken and re-acquired only after it has been dropped.Found this while reviewing #2759, which changes the cleanup function from
fprintf(stderr, ...)toPyErr_WarnExrequireing the gil. Auditing where reports can be emitted from surfaceddeviceptr_import_ipcas the onepw_*call site inside a held C++ lock.How to Reproduce
Reproducer
Four threads each import a distinct descriptor, so all of them take the slow path. Every import succeeds — nothing fails and nothing is reported — which is what shows the hang is purely the lock ordering rather than anything on an error path.
Expected: four successful imports, exit 0.
Actual: the importing process never completes and is killed by its own watchdog.
Expected behavior
No deadlock.
Operating System
No response
nvidia-smi output
No response