Repository navigation
ENH: Reduce refcount contention on heap types in free-threaded builds - #32853
prathamhole14 wants to merge 1 commit into
Conversation
|
@kumaraditya303 is this a sane approach? I guess it is fine at least when subclassing isn't possible (maybe generally, but dunno), but it feels like a hack (i.e. it's just not the job of (Basically, this is probably fine, with a code comment, but it piles on that we are working around something missing in CPythong: Proper heap type immortalization and maybe even |
This seems fine to me in the case of type which are not subclass-able, I think there could be issues if type is base type and subclass is implemented in C extension with a custom tp_dealloc but pure python subclasses would work. |
|
Well, the main question is if this is worthwhile to worry about (in NumPy, I am assuming that CPython will provide us with better solutions eventually). And I am tempted to say: This seems not like a big enough deal to do something that is only safe because we disallow subclasses, which we do here (and I am very sure I guess one could also call the EDIT: Another way to say why it's a weird: |
PR summary
Follow-up to #32502 (see the discussion in #32552).
The types converted in #32502 drop their type reference with
Py_DECREF(type)in their owntp_dealloc. On free-threaded builds, threads that create these objects then contend on the type's refcount, e.g.a.flagsat 8 threads is about 8x slower than before #32502.This PR removes their
tp_deallocso CPython'ssubtype_deallocis used, which on free-threaded 3.14+ drops the type reference with a per-thread refcount. The cleanup moves totp_free. This costs about 2-4 ns per object on GIL builds.Related to: #32747, #32451 and #31913
Free-threading benchmark results
Setup: AMD Ryzen 7 4800H laptop (8 cores, 16 threads), CPython 3.14.7t and 3.14.6 (GIL), release builds without BLAS.
Throughput is million calls per second, summed over all threads.
1. Slowdown from #32502 (3.14t, commit before vs after the merge)
a.flagsa.flagsa.flagsa.flags.writeablenp.require(a, ['A', 'W'])np.broadcast_arrays(a, b)np.unique(small)np.linspace(0, 1, 5)np.delete(a, 1)2. With the fix (3.14t,
mainat f99b1a0 vsmain+ fix)a.flagsa.flagsa.flagsa.flagsa.flags.writeablenp.require(a, ['A', 'W'])np.broadcast_arrays(a, b)np.unique(small)np.linspace(0, 1, 5)np.delete(a, 1)a.shape(control)3. Single-thread cost on a GIL build (3.14, ns per call, 5 runs per build)
a.shape(control, not touched)np.shape(a)(calls a dispatcher, creates none)a.flagsa.flags.writeablenp.require(a, ['A', 'W'])np.broadcast_arrays(a, b)AI Disclosure
To implement the changes and benchmark with claude opus 5.5 max, I have reviewed the code.