Summary
If GraphBuilder.join raises partway through, the forked builders it has not yet closed are left open with capturing streams. The test that hit this failed cleanly, but the process then segfaulted during gc.collect() in the init_cuda fixture teardown while those abandoned objects were destroyed. An exception inside join should not be able to crash the interpreter later.
What was observed
On PR #2750, a transient bug made the temporary ordering event in Stream.wait fail to be created. join calls root_bdr.stream.wait(builder.stream) and then builder.close() for each non-root builder; the wait raised on the first builder, so no forked builder was closed. Every GPU test job then crashed with:
Fatal Python error: Segmentation fault
Current thread ... (most recent call first):
File ".../cuda_core/tests/conftest.py", line 192 in init_cuda
Line 192 is the gc.collect() in the fixture's finally. The first tests to fail were test_graph_complete_after_close_forked and test_graph_definition_raises_for_forked in tests/graph/test_graph_builder.py, both of which go through split and join.
https://github.com/NVIDIA/cuda-python/actions/runs/33911046695/job/101149369210
The event-creation bug is fixed in that PR, so the crash is no longer reachable through this route. The teardown fragility remains: any exception in join (or an interrupted split/join sequence) leaves builders in the same state.
Suggested direction
- Make
join exception-safe: on failure, close or otherwise neutralize the forked builders that were not joined, rather than leaving them mid-capture.
- Make the forked-builder destructor tolerant of an abandoned capture, so destruction of a never-joined fork cannot dereference an invalid handle or end capture on a stream that is no longer valid. Identifying the exact dereference is part of this issue; the Python traceback stops at
gc.collect().
Reproduction sketch
gb = Device().create_graph_builder().begin_building()
left, right = gb.split(2)
# force root_bdr.stream.wait(...) to raise inside join, e.g. by monkeypatching Stream.wait
with pytest.raises(Exception):
GraphBuilder.join(left, right)
del left, right, gb
gc.collect() # crashes today
Summary
If
GraphBuilder.joinraises partway through, the forked builders it has not yet closed are left open with capturing streams. The test that hit this failed cleanly, but the process then segfaulted duringgc.collect()in theinit_cudafixture teardown while those abandoned objects were destroyed. An exception insidejoinshould not be able to crash the interpreter later.What was observed
On PR #2750, a transient bug made the temporary ordering event in
Stream.waitfail to be created.joincallsroot_bdr.stream.wait(builder.stream)and thenbuilder.close()for each non-root builder; thewaitraised on the first builder, so no forked builder was closed. Every GPU test job then crashed with:Line 192 is the
gc.collect()in the fixture'sfinally. The first tests to fail weretest_graph_complete_after_close_forkedandtest_graph_definition_raises_for_forkedintests/graph/test_graph_builder.py, both of which go throughsplitandjoin.https://github.com/NVIDIA/cuda-python/actions/runs/33911046695/job/101149369210
The event-creation bug is fixed in that PR, so the crash is no longer reachable through this route. The teardown fragility remains: any exception in
join(or an interrupted split/join sequence) leaves builders in the same state.Suggested direction
joinexception-safe: on failure, close or otherwise neutralize the forked builders that were not joined, rather than leaving them mid-capture.gc.collect().Reproduction sketch