Visitar URL original
Using latest cothread on python3.11 with pyqt results in segfault · Issue #75 · DiamondLightSource/cothread · GitHub
Skip to content

Using latest cothread on python3.11 with pyqt results in segfault #75

Description

@ptsOSL

Calling this: cothread.iqt(run_exec=True)

Results in a segfault.

UPDATE 2

Im seeing some issues with following the advice in update 1, things not working the way they used too.

UPDATE 1

For the first fix below, QApplication.exec() should be spawned as a cothread, not called directly. Such as:
cothread.Spawn(_qapp.exec)

This works fine, but if you do any future yields or call any cothread functions which yield from the main thread it will segfault.

So this works:

def ticker():
    while True:
        cothread.Sleep(1)
        print("*** --------------------------------------------- tick")

cothread.Spawn(_qapp.exec)
cothread.Spawn(ticker)
cothread.WaitForQuit()

But this segfaults:

cothread.Spawn(_qapp.exec)
cothread.Yield()
cothread.WaitForQuit()

This is okayish as you typically spawning the exec thread is the last thing you do before running cothread.WaitForQuit().
But obviously the root issue needs to be fixed.

The second fix below suffers from the same issue.

Quick fixes:

  • A quick fix can be made by either setting run_exec=False and calling QApplication.exec() manually.

  • Or by removing the cothread.Yield() called immediately after cothread.Spawn(getattr(_qapp, exec_name), stack_size = QT_STACK_SIZE) in the iqt() function in input_hook.py

I don't yet understand the root cause of this issue, but will continue investigating.

Activity

  1. Araneidae commented on Jun 19, 2025

    @Araneidae
    Collaborator

    Can you specify which version of PyQt, please?

    I remember PyQt support being a gnarly and a bit scary (there is some evil under the hood), and I think this code hasn't had much attention in some years. Do let me know what you find, and I think we'll try and reproduce here.

    A complete reproducible example would be great.

  2. ptsOSL commented on Jun 19, 2025

    @ptsOSL
    Author

    I have seen it on both pyqt5 and pyqt6, specifically i am using:
    PyQt5 5.15.11
    PyQt5-Qt5 5.15.17
    PyQt5_sip 12.17.0
    PyQt6 6.9.1
    PyQt6-Qt6 6.9.1
    PyQt6_sip 13.10.2

    Minimum working example is:

    import cothread
    _qapp = cothread.iqt(run_exec=True)
    cothread.WaitForQuit()
    

    Which results in:
    Segmentation fault (core dumped)

  3. ptsOSL commented on Jun 19, 2025

    @ptsOSL
    Author

    The segfault is firing when we attempt to Yield from the main thread after calling cothread.Spawn(_qapp.exec).

    Yield() -> do_yield() -> _Scheduler.wait_until()

    In the wait_until() function we call _coroutine.switch().

    Naturally this attempts to switch coroutines, which involves swapping out stack frames.

    Something is wrong with the stack frame that it switches too, as we segfault as soon as we return to python.

  4. ptsOSL commented on Jun 26, 2025

    @ptsOSL
    Author

    Python3.11:

    • Without pyqt, cothreads appears to work fine
    • With pyqt, we get a segfault

    Python3.11-debug:
    -We hit the following check in pystate.c:

    #if defined(Py_DEBUG)
        if (newts) {
            /* This can be called from PyEval_RestoreThread(). Similar
               to it, we need to ensure errno doesn't change.
            */
            int err = errno;
            PyThreadState *check = _PyGILState_GetThisThreadState(gilstate);
            if (check && check->interp == newts->interp && check != newts)
                Py_FatalError("Invalid thread state for this thread");
            errno = err;
        }
    #endif
    

    -This check was removed shortly after as it was decided that the things it was checking against was okay to do.
    -If this is removed then we fail the first assertion in the following function in ceval.c:

    static void
    _PyEvalFrameClearAndPop(PyThreadState *tstate, _PyInterpreterFrame * frame)
    {
        // Make sure that this is, indeed, the top frame. We can't check this in
        // _PyThreadState_PopFrame, since f_code is already cleared at that point:
        assert((PyObject **)frame + frame->f_code->co_nlocalsplus +
            frame->f_code->co_stacksize + FRAME_SPECIALS_SIZE == tstate->datastack_top);
        tstate->recursion_remaining--;
        assert(frame->frame_obj == NULL || frame->frame_obj->f_frame == frame);
        assert(frame->owner == FRAME_OWNED_BY_THREAD);
        _PyFrame_Clear(frame);
        tstate->recursion_remaining++;
        _PyThreadState_PopFrame(tstate, frame);
    }
    

    -It appears that our datastack_top pointer is out of sync with where python thinks it should be, in testing the values have only been off by a few hundred bytes.

    Python3.12:
    -Works normally for the tests i did, including pyqt, but this may just be "lucky"

    Python3.12-debug:
    -Also seems to work fine

    Python3.13:
    -Initially I was seeing it not switch to the cothreads I made, but I cant recreate this anymore and it seems to be working fine now.

    Python3.13-debug:
    -Also seems to work fine

  5. AlexanderWells-diamond commented on Jun 27, 2025

    @AlexanderWells-diamond
    Contributor

    I'm having to stop work on this for the moment, so this message is a writeup of all I've learned so far.

    First, the original issue @ptsOSL found. It appears to be a segfault that happens due to a malformed stack or somesuch - simply stepping down/up of the stack in a debugger at a certain point causes the segfault. Nothing obvious in the debugger itself shows an issue, so its likely the stack is somehow corrupted. This is repeatable and reliably occurs.

    Initial investigation into it was done by building Python --with-pydebug, which enables extra checks. The first error we see is:

    Fatal Python error: _PyThreadState_Swap: Invalid thread state for this thread
    

    This appears to be a check that should not exist, and was removed in Python3.12 in this PR. Doing a manual Python build that removes this check gives us a new error message:

    Fatal Python error: _PyMem_DebugMalloc: Python memory allocator called without holding the GIL
    

    So it appears that our use of the various ThreadState_* APIs is incorrect.

    My first attempt to solve this was to try to mimic greenlet as it does something very similar to cothread (And Michael commented an earlier version of cothread was based off of greenlet). Greenlet seems to be storing and restoring the threadstate objects here and here respectively. There is a branch of this code available here.

    The code that actually does the resetting of the threadstate is still a mystery to me. I think it is done by this save statement which is then read by macros used in the ASM, but I can't be sure.

    Anyway, that branch doesn't work and results in this error message:

    Python/ceval.c:4780: _Py_NegativeRefcount: Assertion failed: object has negative ref count
    Memory block allocated at (most recent call first):
      File "<frozen importlib._bootstrap_external>", line 1241
    
    object address  : 0x7fe9168b12c0
    object refcount : -1
    object type     : 0xa383c0
    object type name: tuple
    object repr     : <refcnt -1 at 0x7fe9168b12c0>
    

    The object with the negative reference count appears to be something on the stack of the Python function being called - but I have been unable to print it in any useful format (attempting to call the repr function via the C API just results in more errors).

    I asked a question on the greenlet repo in case they can offer any help; that can be found here

    The second attempt was to mess around with the GIL APIs to try and satisfy the original error that we don't hold the GIL. That branch can be seen here. Broadly speaking this seems to accomplish its goal of correctly releasing the GIL in the original ThreadState, then acquiring it in the new one. We do however encounter a different problem later on:

    python: Python/ceval.c:6402: _PyEvalFrameClearAndPop: Assertion `(PyObject **)frame + frame->f_code->co_nlocalsplus + frame->f_code->co_stacksize + FRAME_SPECIALS_SIZE == tstate->datastack_top' failed.
    Fatal Python error: Aborted
    

    This error can be found here. It was only added to Python3.11 and did not appear in 3.10. This was done in this PR. It appears to be trying to ensure that the ThreadState and the CFrame information are both pointing to the same thing, which for some reason we are not doing.

    I asked for help on this topic on the Python help forums

    And finally, here's a list of useful links and documentation I've been reading while researching this topic:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions