Bug report
Since 3.14, the warnings module protects its state with a lock in the interpreter, added in #128386 for #128384. PyOS_AfterFork_Child() does not reset this lock. If another thread holds it when a thread calls os.fork(), the child inherits it locked, and the first warning in the child blocks forever. The child does not need to run any code for this to happen. PyOS_AfterFork_Child() clears the states of the other threads. Objects held only in those states, such as threading.local() data, are freed there, and a finalizer can issue a warning before os.fork() returns.
In the reproducer, one thread keeps an unclosed socket in threading.local() and issues warnings in a loop. The child only calls os._exit(0), so a child that hangs is stuck inside os.fork():
import os, socket, sys, threading, time, warnings
# Every warning is shown, so warn() runs showwarning(), which is Python code,
# while it holds the warnings lock.
warnings.simplefilter("always")
warnings.showwarning = lambda *args, **kwargs: None
local = threading.local()
def warn_forever():
local.sock = socket.socket() # never closed: freeing it issues a ResourceWarning
while True:
warnings.warn("from another thread")
threading.Thread(target=warn_forever, daemon=True).start()
time.sleep(0.5)
hung = 0
for _ in range(20):
pid = os.fork()
if pid == 0:
os._exit(0) # the child runs no other code
time.sleep(0.5)
if os.waitpid(pid, os.WNOHANG) == (0, 0):
hung += 1
os.kill(pid, 9)
os.waitpid(pid, 0)
print(f"{sys.version.split()[0]}: {hung} of 20 children hung", flush=True)
os._exit(0) # skip interpreter shutdown, which can hang too (see below)
Results on Linux x86_64 with glibc 2.43:
| Python |
Build |
Children that hung, of 20 |
| 3.12.13 |
built from source |
0 |
| 3.13.16 |
conda-forge |
0 |
| 3.13.16, free-threaded |
conda-forge |
0 |
| 3.14.4 |
Ubuntu package |
17 |
| 3.14.8 |
conda-forge |
20 |
| 3.14.8, free-threaded |
conda-forge |
18 |
| 3.15.0rc3 |
conda-forge |
15 |
gdb on a hung child (3.14.4, Ubuntu build without debug symbols; frames without a name are cut) shows a finalizer that runs in PyThreadState_Clear() and waits for the lock in PyErr_ResourceWarning():
#10 0x00000000005cbfc6 in _PyRecursiveMutex_Lock ()
...
#14 0x000000000047da7c in PyErr_ResourceWarning ()
...
#16 0x00000000005cd2d8 in PyObject_CallFinalizerFromDealloc ()
...
#23 0x00000000005d9bf6 in PyObject_ClearWeakRefs ()
...
#26 0x00000000006d0f22 in PyThreadState_Clear ()
#27 0x00000000004850c2 in PyOS_AfterFork_Child ()
The lock can be held while other threads run. do_warn() holds it across warn_explicit(), and warn_explicit() runs the Python function showwarning() that writes the message:
|
warnings_lock(tstate->interp); |
|
res = warn_explicit(tstate, category, message, filename, lineno, module, registry, |
|
NULL, source); |
|
warnings_unlock(tstate->interp); |
catch_warnings.__enter__() and __exit__(), filterwarnings() and simplefilter() also run Python code while they hold the lock. Every C warning call takes the lock before it checks the filters, so the ResourceWarning in the child blocks although the default filters ignore it.
The lock is not in LOCKS_INIT, the list of runtime locks that the child resets. On main, PyOS_AfterFork_Child() also resets _Py_jit_debug_mutex. It then clears the dead thread states, and the comment above that call notes that this may call destructors:
|
#if defined(PY_HAVE_JIT_GDB_UNWIND) |
|
// The child can inherit this mutex locked if another thread held it at |
|
// fork(), but the child itself cannot be inside gdb_jit_register_code(). |
|
// Reinitialize it before any executor cleanup can unregister JIT code. |
|
_Py_jit_debug_mutex = (PyMutex){0}; |
|
#endif |
|
|
|
reset_remotedebug_data(tstate); |
|
|
|
reset_asyncio_state((_PyThreadStateImpl *)tstate); |
|
|
|
// Remove the dead thread states. We "start the world" once we are the only |
|
// thread state left to undo the stop the world call in `PyOS_BeforeFork`. |
|
// That needs to happen before `_PyThreadState_DeleteList`, because that |
|
// may call destructors. |
|
PyThreadState *list = _PyThreadState_RemoveExcept(tstate); |
|
_PyEval_StartTheWorldAll(&_PyRuntime); |
|
_PyThreadState_DeleteList(list, /*is_after_fork=*/1); |
Holding the lock across the fork avoids the hang. With these two lines added to the reproducer, 0 of 20 children hang on 3.14.4, 3.14.8 and 3.15.0rc3:
import _warnings
os.register_at_fork(before=_warnings._acquire_lock, after_in_parent=_warnings._release_lock, after_in_child=_warnings._release_lock)
A fix could take the lock in PyOS_BeforeFork() and release it after the fork, as is done for the import lock, or reset it in PyOS_AfterFork_Child(), as is done for _Py_jit_debug_mutex.
This hangs JupyterLab terminals. terminado starts them with ptyprocess on the server's event loop thread, and ptyprocess waits for exec() with a blocking read after pty.fork(). When the child hangs inside os.forkpty(), the whole Jupyter server stops responding. IPython runs ! commands through pexpect, and pexpect uses the same code.
I know this is grey territory. The os.fork() documentation says that "it has never been safe to mix threading with os.fork() on POSIX platforms", and earlier reports of hangs in a fork child were closed on that basis (#111635, #112788). Here the child runs no code of its own, and it waits for a lock that CPython added in 3.14. Two earlier issues about a lock or state that the child inherits were fixed:
The same lock can also hang interpreter shutdown, with no fork. At shutdown, the daemon thread is frozen while it holds the lock, and the ResourceWarning for its socket then blocks the main thread. The reproducer without the fork loop and the final os._exit(0) hangs at exit in 10 of 10 runs on 3.14.4, 8 of 10 on 3.14.8 and 9 of 10 on 3.15.0rc3. It hangs in 0 of 10 runs on 3.13.16, and on 3.14.4 without the socket. I can open a separate issue for it.
I used Opus 5.5 for assistance with investigation and the writeup.
CPython versions tested on:
3.12, 3.13, 3.14, 3.15
Operating systems tested on:
Linux
Linked PRs
Bug report
Since 3.14, the warnings module protects its state with a lock in the interpreter, added in #128386 for #128384.
PyOS_AfterFork_Child()does not reset this lock. If another thread holds it when a thread callsos.fork(), the child inherits it locked, and the first warning in the child blocks forever. The child does not need to run any code for this to happen.PyOS_AfterFork_Child()clears the states of the other threads. Objects held only in those states, such asthreading.local()data, are freed there, and a finalizer can issue a warning beforeos.fork()returns.In the reproducer, one thread keeps an unclosed socket in
threading.local()and issues warnings in a loop. The child only callsos._exit(0), so a child that hangs is stuck insideos.fork():Results on Linux x86_64 with glibc 2.43:
gdb on a hung child (3.14.4, Ubuntu build without debug symbols; frames without a name are cut) shows a finalizer that runs in
PyThreadState_Clear()and waits for the lock inPyErr_ResourceWarning():The lock can be held while other threads run.
do_warn()holds it acrosswarn_explicit(), andwarn_explicit()runs the Python functionshowwarning()that writes the message:cpython/Python/_warnings.c
Lines 1136 to 1139 in 1610807
catch_warnings.__enter__()and__exit__(),filterwarnings()andsimplefilter()also run Python code while they hold the lock. Every C warning call takes the lock before it checks the filters, so theResourceWarningin the child blocks although the default filters ignore it.The lock is not in
LOCKS_INIT, the list of runtime locks that the child resets. On main,PyOS_AfterFork_Child()also resets_Py_jit_debug_mutex. It then clears the dead thread states, and the comment above that call notes that this may call destructors:cpython/Modules/posixmodule.c
Lines 782 to 799 in 1610807
Holding the lock across the fork avoids the hang. With these two lines added to the reproducer, 0 of 20 children hang on 3.14.4, 3.14.8 and 3.15.0rc3:
A fix could take the lock in
PyOS_BeforeFork()and release it after the fork, as is done for the import lock, or reset it inPyOS_AfterFork_Child(), as is done for_Py_jit_debug_mutex.This hangs JupyterLab terminals. terminado starts them with ptyprocess on the server's event loop thread, and ptyprocess waits for
exec()with a blocking read afterpty.fork(). When the child hangs insideos.forkpty(), the whole Jupyter server stops responding. IPython runs!commands through pexpect, and pexpect uses the same code.I know this is grey territory. The
os.fork()documentation says that "it has never been safe to mix threading with os.fork() on POSIX platforms", and earlier reports of hangs in a fork child were closed on that basis (#111635, #112788). Here the child runs no code of its own, and it waits for a lock that CPython added in 3.14. Two earlier issues about a lock or state that the child inherits were fixed:HEAD_LOCKlocked and deadlocked on it inPyOS_AfterFork_Child(). The fix was backported to 3.8.reset_asyncio_state()toPyOS_AfterFork_Child().The same lock can also hang interpreter shutdown, with no fork. At shutdown, the daemon thread is frozen while it holds the lock, and the
ResourceWarningfor its socket then blocks the main thread. The reproducer without the fork loop and the finalos._exit(0)hangs at exit in 10 of 10 runs on 3.14.4, 8 of 10 on 3.14.8 and 9 of 10 on 3.15.0rc3. It hangs in 0 of 10 runs on 3.13.16, and on 3.14.4 without the socket. I can open a separate issue for it.I used Opus 5.5 for assistance with investigation and the writeup.
CPython versions tested on:
3.12, 3.13, 3.14, 3.15
Operating systems tested on:
Linux
Linked PRs