Repository navigation
Using latest cothread on python3.11 with pyqt results in segfault #75
Description
Activity
Can you specify which version of PyQt, please?
I remember PyQt support being a gnarly and a bit scary (there is some evil under the hood), and I think this code hasn't had much attention in some years. Do let me know what you find, and I think we'll try and reproduce here.
A complete reproducible example would be great.
I have seen it on both pyqt5 and pyqt6, specifically i am using:
PyQt5 5.15.11
PyQt5-Qt5 5.15.17
PyQt5_sip 12.17.0
PyQt6 6.9.1
PyQt6-Qt6 6.9.1
PyQt6_sip 13.10.2Minimum working example is:
import cothread _qapp = cothread.iqt(run_exec=True) cothread.WaitForQuit()Which results in:
Segmentation fault (core dumped)The segfault is firing when we attempt to Yield from the main thread after calling cothread.Spawn(_qapp.exec).
Yield() -> do_yield() -> _Scheduler.wait_until()
In the wait_until() function we call _coroutine.switch().
Naturally this attempts to switch coroutines, which involves swapping out stack frames.
Something is wrong with the stack frame that it switches too, as we segfault as soon as we return to python.
Python3.11:
- Without pyqt, cothreads appears to work fine
- With pyqt, we get a segfault
Python3.11-debug:
-We hit the following check in pystate.c:#if defined(Py_DEBUG) if (newts) { /* This can be called from PyEval_RestoreThread(). Similar to it, we need to ensure errno doesn't change. */ int err = errno; PyThreadState *check = _PyGILState_GetThisThreadState(gilstate); if (check && check->interp == newts->interp && check != newts) Py_FatalError("Invalid thread state for this thread"); errno = err; } #endif-This check was removed shortly after as it was decided that the things it was checking against was okay to do.
-If this is removed then we fail the first assertion in the following function in ceval.c:static void _PyEvalFrameClearAndPop(PyThreadState *tstate, _PyInterpreterFrame * frame) { // Make sure that this is, indeed, the top frame. We can't check this in // _PyThreadState_PopFrame, since f_code is already cleared at that point: assert((PyObject **)frame + frame->f_code->co_nlocalsplus + frame->f_code->co_stacksize + FRAME_SPECIALS_SIZE == tstate->datastack_top); tstate->recursion_remaining--; assert(frame->frame_obj == NULL || frame->frame_obj->f_frame == frame); assert(frame->owner == FRAME_OWNED_BY_THREAD); _PyFrame_Clear(frame); tstate->recursion_remaining++; _PyThreadState_PopFrame(tstate, frame); }-It appears that our datastack_top pointer is out of sync with where python thinks it should be, in testing the values have only been off by a few hundred bytes.
Python3.12:
-Works normally for the tests i did, including pyqt, but this may just be "lucky"Python3.12-debug:
-Also seems to work finePython3.13:
-Initially I was seeing it not switch to the cothreads I made, but I cant recreate this anymore and it seems to be working fine now.Python3.13-debug:
-Also seems to work fine- added a commit that references this issue
on Jun 27, 2025 AlexanderWells-diamond commented
on Jun 27, 2025 ContributorMore actionsI'm having to stop work on this for the moment, so this message is a writeup of all I've learned so far.
First, the original issue @ptsOSL found. It appears to be a segfault that happens due to a malformed stack or somesuch - simply stepping down/up of the stack in a debugger at a certain point causes the segfault. Nothing obvious in the debugger itself shows an issue, so its likely the stack is somehow corrupted. This is repeatable and reliably occurs.
Initial investigation into it was done by building Python
--with-pydebug, which enables extra checks. The first error we see is:Fatal Python error: _PyThreadState_Swap: Invalid thread state for this threadThis appears to be a check that should not exist, and was removed in Python3.12 in this PR. Doing a manual Python build that removes this check gives us a new error message:
Fatal Python error: _PyMem_DebugMalloc: Python memory allocator called without holding the GILSo it appears that our use of the various
ThreadState_*APIs is incorrect.My first attempt to solve this was to try to mimic
greenletas it does something very similar tocothread(And Michael commented an earlier version of cothread was based off of greenlet). Greenlet seems to be storing and restoring the threadstate objects here and here respectively. There is a branch of this code available here.The code that actually does the resetting of the threadstate is still a mystery to me. I think it is done by this save statement which is then read by macros used in the ASM, but I can't be sure.
Anyway, that branch doesn't work and results in this error message:
Python/ceval.c:4780: _Py_NegativeRefcount: Assertion failed: object has negative ref count Memory block allocated at (most recent call first): File "<frozen importlib._bootstrap_external>", line 1241 object address : 0x7fe9168b12c0 object refcount : -1 object type : 0xa383c0 object type name: tuple object repr : <refcnt -1 at 0x7fe9168b12c0>The object with the negative reference count appears to be something on the stack of the Python function being called - but I have been unable to print it in any useful format (attempting to call the repr function via the C API just results in more errors).
I asked a question on the greenlet repo in case they can offer any help; that can be found here
The second attempt was to mess around with the GIL APIs to try and satisfy the original error that we don't hold the GIL. That branch can be seen here. Broadly speaking this seems to accomplish its goal of correctly releasing the GIL in the original ThreadState, then acquiring it in the new one. We do however encounter a different problem later on:
python: Python/ceval.c:6402: _PyEvalFrameClearAndPop: Assertion `(PyObject **)frame + frame->f_code->co_nlocalsplus + frame->f_code->co_stacksize + FRAME_SPECIALS_SIZE == tstate->datastack_top' failed. Fatal Python error: AbortedThis error can be found here. It was only added to Python3.11 and did not appear in 3.10. This was done in this PR. It appears to be trying to ensure that the ThreadState and the CFrame information are both pointing to the same thing, which for some reason we are not doing.
I asked for help on this topic on the Python help forums
And finally, here's a list of useful links and documentation I've been reading while researching this topic:
- The official docs on the C API: https://docs.python.org/3/c-api/init.html#thread-state-and-the-global-interpreter-lock
- The changelog for Python 3.11 noting the
framedisappearence: https://docs.python.org/3/whatsnew/3.11.html#whatsnew311-c-api-porting - Someone from
pyodidedoing a very similar thing and copyinggreenlet's thread manipulation (unfortunately some important links are dead): https://discuss.python.org/t/api-for-stack-switching-save-state-restore-state/26390 - Some WIP of adding exactly what we want to cpython itself: Add "unstable" frame stack api python/cpython#91371
- The original issue that added the assertion we're encountering when using the GIL manipulation method: Python/pystate.c:2218: _PyThreadState_PopFrame: Assertion `tstate->datastack_top >= base' failed. python/cpython#93252
- pyoide's version of thread state swapping: https://github.com/pyodide/pyodide/blob/d317e66d174efa203c59f094eb7714c0b90f7613/src/core/stack_switching/pystate.c#L222
- A very promising issue report that may contain further details on how Greenlet works and what we must do: https://bugs.python.org/issue46090
Calling this: cothread.iqt(run_exec=True)
Results in a segfault.
UPDATE 2
UPDATE 1
Quick fixes:
A quick fix can be made by either setting run_exec=False and calling QApplication.exec() manually.
Or by removing the
cothread.Yield()called immediately aftercothread.Spawn(getattr(_qapp, exec_name), stack_size = QT_STACK_SIZE)in the iqt() function in input_hook.pyI don't yet understand the root cause of this issue, but will continue investigating.