Repository navigation
Potential deadlock in test_tracemalloc with the FT build and tail-call dispatch. #139116
Description
Activity
- changed the title
[-]Possibility of deadlock at test_tracmalloc at FT build with tail-call dispatch.[/-][+]Possibility of deadlock o test_tracmalloc with FT build and tail-call dispatch.[/+]on Sep 18, 2025 - changed the title
[-]Possibility of deadlock o test_tracmalloc with FT build and tail-call dispatch.[/-][+]Possibility of deadlock at test_tracmalloc with FT build and tail-call dispatch.[/+]on Sep 18, 2025 - changed the title
[-]Possibility of deadlock at test_tracmalloc with FT build and tail-call dispatch.[/-][+]Potential deadlock in test_tracemalloc with the FT build and tail-call dispatch.[/+]on Sep 18, 2025 Probably a FT bug I doubt it's anything to do with TC.
Reacted by Donghee Na- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Sep 18, 2025 Happens frequently on Windows, too (without tail-calling):
https://buildbot.python.org/api/v2/logs/14575096/raw_inline
timeout after 20 min.Successful re-run takes less than a second.
ASAN makes this easier to reproduce: https://buildbot.python.org/#/builders/1366
This started after #137994 was merged. The deadlock is in:
- One thread in
_PyTraceMalloc_Stop, withTABLES_LOCKheld, callingPyRefTracer_SetTracerwhich wants to stop the world - Another thread in
PyTraceMalloc_Track, just attached thread state, waiting forTABLES_LOCK
GDB backtraces (click to open)
Main thread waiting to stop the world in
_PyTraceMalloc_Stop:(gdb) bt #0 __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S:56 #1 0x00007ffff758e75c in __internal_syscall_cancel (a1=<optimized out>, a2=<optimized out>, a3=a3@entry=0, a4=<optimized out>, a5=a5@entry=0, a6=a6@entry=4294967295, nr=202) at cancellation.c:49 #2 0x00007ffff758edcc in __futex_abstimed_wait_common64 (private=<optimized out>, futex_word=0x7bfff5202100, expected=0, op=<optimized out>, abstime=0x7bfff50a8040, cancel=true) at futex-internal.c:57 #3 __futex_abstimed_wait_common (futex_word=futex_word@entry=0x7bfff5202100, expected=expected@entry=0, clockid=clockid@entry=1, abstime=abstime@entry=0x7bfff50a8040, private=<optimized out>, cancel=cancel@entry=true) at futex-internal.c:87 #4 0x00007ffff758ee2f in __GI___futex_abstimed_wait_cancelable64 (futex_word=futex_word@entry=0x7bfff5202100, expected=expected@entry=0, clockid=clockid@entry=1, abstime=abstime@entry=0x7bfff50a8040, private=<optimized out>) at futex-internal.c:139 #5 0x00007ffff75995b0 in do_futex_wait (sem=sem@entry=0x7bfff5202100, clockid=clockid@entry=1, abstime=abstime@entry=0x7bfff50a8040) at /usr/src/debug/glibc-2.41-11.fc42.x86_64/nptl/sem_waitcommon.c:111 #6 0x00007ffff7599656 in __new_sem_wait_slow64 (sem=0x7bfff5202100, clockid=1, abstime=0x7bfff50a8040) at /usr/src/debug/glibc-2.41-11.fc42.x86_64/nptl/sem_waitcommon.c:183 #7 0x00007ffff75996d5 in ___sem_clockwait64 (sem=sem@entry=0x7bfff5202100, clockid=clockid@entry=1, abstime=abstime@entry=0x7bfff50a8040) at sem_clockwait.c:46 #8 0x0000000000aa847d in _PySemaphore_PlatformWait (sema=0x7bfff5202100, timeout=1000000) at Python/parking_lot.c:159 #9 _PySemaphore_Wait (sema=<optimized out>, timeout=<optimized out>, detach=<optimized out>) at Python/parking_lot.c:242 #10 0x0000000000aa8d4a in _PyParkingLot_Park (addr=addr@entry=0x1139564 <_PyRuntime+10532>, expected=expected@entry=0x7bfff50a7fc0, size=size@entry=1, timeout_ns=<optimized out>, park_arg=park_arg@entry=0x0, detach=<optimized out>) at Python/parking_lot.c:345 #11 0x0000000000a8c2ab in PyEvent_WaitTimed (evt=evt@entry=0x1139564 <_PyRuntime+10532>, timeout_ns=timeout_ns@entry=1000000, detach=detach@entry=0) at Python/lock.c:309 #12 0x0000000000ac393f in stop_the_world (stw=0x1139560 <_PyRuntime+10528>) at Python/pystate.c:2296 #13 0x0000000000acbf2c in _PyEval_StopTheWorldAll (runtime=<optimized out>) at Python/pystate.c:2342 #14 0x00000000006a373d in PyRefTracer_SetTracer (tracer=tracer@entry=0x0, data=data@entry=0x0) at Objects/object.c:3291 #15 0x0000000000b30b0c in _PyTraceMalloc_Stop () at Python/tracemalloc.c:873 #16 0x0000000000bd8799 in _tracemalloc_stop_impl (module=<optimized out>) at ./Modules/_tracemalloc.c:118 #17 _tracemalloc_stop (module=<optimized out>, _unused_ignored=<optimized out>) at ./Modules/clinic/_tracemalloc.c.h:134 #18 0x0000000000538004 in _PyObject_VectorcallTstate (tstate=0x118beb8 <_PyRuntime+348792>, callable=<optimized out>, args=0x0, nargsf=0, kwnames=0x0) at ./Include/internal/pycore_call.h:169 #19 PyObject_CallNoArgs (func=<built-in method stop of module object at remote 0x7bffa74ed8e0>) at Objects/call.c:106 #20 0x00007bfff4d4b810 in tracemalloc_track_race (self=<optimized out>, args=<optimized out>) at ./Modules/_testcapi/mem.c:650 #21 0x000000000053926d in _PyObject_VectorcallTstate (tstate=0x118beb8 <_PyRuntime+348792>, callable=<optimized out>, args=<optimized out>, nargsf=<optimized out>, kwnames=<optimized out>) at ./Include/internal/pycore_call.h:169 #22 PyObject_Vectorcall (callable=<built-in method tracemalloc_track_race of module object at remote 0x7bffa6f361d0>, args=<optimized out>, nargsf=<optimized out>, kwnames=<optimized out>) at Objects/call.c:327 #23 0x000000000093d023 in _PyEval_EvalFrameDefault (tstate=<optimized out>, frame=<optimized out>, throwflag=<optimized out>) at Python/generated_cases.c.h:1620 #24 0x0000000000961516 in _PyEval_EvalFrame (tstate=0x118beb8 <_PyRuntime+348792>, frame=0x7e8ff63e6360, throwflag=0) at ./Include/internal/pycore_ceval.h:121 #25 _PyEval_Vector (tstate=<optimized out>, func=<optimized out>, locals=0x0, args=<optimized out>, argcount=<optimized out>, kwnames=<optimized out>) --Type <RET> for more, q to quit, c to continue without paging--q QuitThread in
TABLES_LOCKinPyTraceMalloc_Track:(gdb) thread 50 [Switching to thread 50 (Thread 0x7bffa2ef26c0 (LWP 247836))] #0 __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S:56 56 ret (gdb) bt #0 __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S:56 #1 0x00007ffff758e75c in __internal_syscall_cancel (a1=<optimized out>, a2=<optimized out>, a3=a3@entry=0, a4=<optimized out>, a5=a5@entry=0, a6=a6@entry=4294967295, nr=202) at cancellation.c:49 #2 0x00007ffff758edcc in __futex_abstimed_wait_common64 (private=<optimized out>, futex_word=0x7bffa03e7100, expected=0, op=<optimized out>, abstime=0x0, cancel=true) at futex-internal.c:57 #3 __futex_abstimed_wait_common (futex_word=futex_word@entry=0x7bffa03e7100, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0, private=<optimized out>, cancel=cancel@entry=true) at futex-internal.c:87 #4 0x00007ffff758ee2f in __GI___futex_abstimed_wait_cancelable64 (futex_word=futex_word@entry=0x7bffa03e7100, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0, private=<optimized out>) at futex-internal.c:139 #5 0x00007ffff759a2cf in do_futex_wait (sem=sem@entry=0x7bffa03e7100, abstime=0x0, clockid=0) at /usr/src/debug/glibc-2.41-11.fc42.x86_64/nptl/sem_waitcommon.c:111 #6 0x00007ffff759a368 in __new_sem_wait_slow64 (sem=0x7bffa03e7100, abstime=0x0, clockid=0) at /usr/src/debug/glibc-2.41-11.fc42.x86_64/nptl/sem_waitcommon.c:183 #7 0x0000000000aa85bc in _PySemaphore_PlatformWait (sema=0x7bffa03e7100, timeout=<optimized out>) at Python/parking_lot.c:172 #8 _PySemaphore_Wait (sema=<optimized out>, timeout=<optimized out>, detach=<optimized out>) at Python/parking_lot.c:242 #9 0x0000000000aa8d4a in _PyParkingLot_Park (addr=addr@entry=0x11394e8 <_PyRuntime+10408>, expected=expected@entry=0x7bffa01e86b0, size=size@entry=1, timeout_ns=timeout_ns@entry=-1, park_arg=park_arg@entry=0x7bffa01e86e0, detach=detach@entry=0) at Python/parking_lot.c:345 #10 0x0000000000a8b2e9 in _PyMutex_LockTimed (m=0x11394e8 <_PyRuntime+10408>, timeout=timeout@entry=-1, flags=flags@entry=_Py_LOCK_DONT_DETACH) at Python/lock.c:120 #11 0x0000000000b311d3 in PyMutex_LockFlags (m=<optimized out>, flags=_Py_LOCK_DONT_DETACH) at ./Include/internal/pycore_lock.h:65 #12 PyTraceMalloc_Track (domain=domain@entry=123, ptr=ptr@entry=10, size=size@entry=1) at Python/tracemalloc.c:1201 #13 0x00007bfff4d48d80 in tracemalloc_track_race_thread (data=data@entry=0x7c1ff63e7c30) at ./Modules/_testcapi/mem.c:591 #14 0x0000000000b213e9 in pythread_wrapper (arg=arg@entry=0x7c1ff63e7c70) at Python/thread_pthread.h:234 #15 0x00007ffff7828ee6 in asan_thread_start (arg=0x7ffff6750000) at ../../../../libsanitizer/asan/asan_interceptors.cpp:239 #16 0x00007ffff7591f54 in start_thread (arg=<optimized out>) at pthread_create.c:448 #17 0x00007ffff761532c in __GI___clone3 () at ../sysdeps/unix/sysv/linux/x86_64/clone3.S:78Another thread waiting for the GIL in
PyTraceMalloc_Track(looks unrelated):(gdb) thread 51 [Switching to thread 51 (Thread 0x7bffa49fd6c0 (LWP 247837))] #0 __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S:56 56 ret (gdb) bt #0 __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S:56 #1 0x00007ffff758e75c in __internal_syscall_cancel (a1=<optimized out>, a2=<optimized out>, a3=a3@entry=0, a4=<optimized out>, a5=a5@entry=0, a6=a6@entry=4294967295, nr=202) at cancellation.c:49 #2 0x00007ffff758edcc in __futex_abstimed_wait_common64 (private=<optimized out>, futex_word=0x7bffa39fd100, expected=0, op=<optimized out>, abstime=0x0, cancel=true) at futex-internal.c:57 #3 __futex_abstimed_wait_common (futex_word=futex_word@entry=0x7bffa39fd100, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0, private=<optimized out>, cancel=cancel@entry=true) at futex-internal.c:87 #4 0x00007ffff758ee2f in __GI___futex_abstimed_wait_cancelable64 (futex_word=futex_word@entry=0x7bffa39fd100, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0, private=<optimized out>) at futex-internal.c:139 #5 0x00007ffff759a2cf in do_futex_wait (sem=sem@entry=0x7bffa39fd100, abstime=0x0, clockid=0) at /usr/src/debug/glibc-2.41-11.fc42.x86_64/nptl/sem_waitcommon.c:111 #6 0x00007ffff759a368 in __new_sem_wait_slow64 (sem=0x7bffa39fd100, abstime=0x0, clockid=0) at /usr/src/debug/glibc-2.41-11.fc42.x86_64/nptl/sem_waitcommon.c:183 #7 0x0000000000aa85bc in _PySemaphore_PlatformWait (sema=0x7bffa39fd100, timeout=<optimized out>) at Python/parking_lot.c:172 #8 _PySemaphore_Wait (sema=<optimized out>, timeout=<optimized out>, detach=<optimized out>) at Python/parking_lot.c:242 #9 0x0000000000aa8d4a in _PyParkingLot_Park (addr=addr@entry=0x7e7ff657012c, expected=expected@entry=0x7bffa37fd640, size=size@entry=4, timeout_ns=timeout_ns@entry=-1, park_arg=park_arg@entry=0x0, detach=detach@entry=0) at Python/parking_lot.c:345 #10 0x0000000000ace553 in tstate_wait_attach (tstate=<optimized out>) at Python/pystate.c:2058 #11 _PyThreadState_Attach (tstate=<optimized out>) at Python/pystate.c:2096 #12 _PyThreadState_Attach (tstate=tstate@entry=0x7e7ff6570100) at Python/pystate.c:2073 #13 0x0000000000a270da in PyEval_RestoreThread (tstate=tstate@entry=0x7e7ff6570100) at Python/ceval_gil.c:654 #14 0x0000000000acd775 in PyGILState_Ensure () at Python/pystate.c:2790 #15 0x0000000000b31158 in PyTraceMalloc_Track (domain=domain@entry=123, ptr=ptr@entry=10, size=size@entry=1) at Python/tracemalloc.c:1200 #16 0x00007bfff4d48d80 in tracemalloc_track_race_thread (data=data@entry=0x7c1ff63e7cb0) at ./Modules/_testcapi/mem.c:591 #17 0x0000000000b213e9 in pythread_wrapper (arg=arg@entry=0x7c1ff63e7cf0) at Python/thread_pthread.h:234 #18 0x00007ffff7828ee6 in asan_thread_start (arg=0x7ffff6708000) at ../../../../libsanitizer/asan/asan_interceptors.cpp:239 #19 0x00007ffff7591f54 in start_thread (arg=<optimized out>) at pthread_create.c:448 #20 0x00007ffff761532c in __GI___clone3 () at ../sysdeps/unix/sysv/linux/x86_64/clone3.S:78- One thread in
Will check this out but in the interim we can revert that commit
It seems to me that a revert would trade one instance of thread unsafety for another. For the interim it might be better to add
asan=Trueto that test's@skip_if_sanitizer.also cc @colesbury -- you might have a better feel for how the locks should be structured here
- added a commit that references this issue
on Sep 30, 2025 Why doesn't
TABLES_LOCKdetach the thread state while waiting in the first place?Why doesn't TABLES_LOCK detach the thread state while waiting in the first place?
Good question.
TABLES_LOCK()uses the_Py_LOCK_DONT_DETACHflag:#define TABLES_LOCK() PyMutex_LockFlags(&tables_lock, _Py_LOCK_DONT_DETACH) #define TABLES_UNLOCK() PyMutex_Unlock(&tables_lock)
Why doesn't TABLES_LOCK detach the thread state while waiting in the first place?
Because you said here:
I don't think it matters.
TABLES_LOCK()didn't detach the thread state before, no need to now. (In fact, some unpredictable things might happen if another thread takes the GIL and then uses tracemalloc.)That was before registering a tracing function started stopping the world, of course.
If it still doesn't matter,
_PY_LOCK_DETACHlooks like the right fix! AFAICS you can use it even without an attached thread state, for the calls fromPyMem_RawFree.I think I was speculating. Things should work correctly even if the thread state detaches.
Reacted by Petr ViktorinBecause you said here:
I now realize this might be read as unfriendly, so I'd like to correct the tone.
I regard you as a locking expert after all the work you've done in this area; the answer to “Why doesn't TABLES_LOCK detach the thread state” is, unironically, because back then, the best of us didn't see this issue coming :)Reacted by Peter Bierma and Gregory P. Smith- moved this from Todo to Done in Release and Deferred blockers 🚫
on Sep 30, 2025 Sadly, this change introduced another race condition: #144763.
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
See: https://github.com/python/cpython/actions/runs/17826347270/job/50680351302?pr=137800
It's hanged over a hour in CI.
cc @Fidget-Spinner @colesbury
Linked PRs