Visitar URL original
ENH: Reduce refcount contention on heap types in free-threaded builds by prathamhole14 · Pull Request #32853 · numpy/numpy · GitHub
Skip to content

ENH: Reduce refcount contention on heap types in free-threaded builds - #32853

Open
prathamhole14 wants to merge 1 commit into
numpy:mainfrom
prathamhole14:heap_types_ft_dealloc
Open

prathamhole14 wants to merge 1 commit into
numpy:mainfrom
prathamhole14:heap_types_ft_dealloc

Conversation

@prathamhole14

Copy link
Copy Markdown
Member

PR summary

Follow-up to #32502 (see the discussion in #32552).

The types converted in #32502 drop their type reference with Py_DECREF(type) in their own tp_dealloc. On free-threaded builds, threads that create these objects then contend on the type's refcount, e.g. a.flags at 8 threads is about 8x slower than before #32502.

This PR removes their tp_dealloc so CPython's subtype_dealloc is used, which on free-threaded 3.14+ drops the type reference with a per-thread refcount. The cleanup moves to tp_free. This costs about 2-4 ns per object on GIL builds.

Related to: #32747, #32451 and #31913

Free-threading benchmark results

Setup: AMD Ryzen 7 4800H laptop (8 cores, 16 threads), CPython 3.14.7t and 3.14.6 (GIL), release builds without BLAS.

Throughput is million calls per second, summed over all threads.

1. Slowdown from #32502 (3.14t, commit before vs after the merge)

Workload Threads Before After After / before
a.flags 1 24.9 18.8 0.76
a.flags 2 49.6 11.2 0.23
a.flags 8 152.6 19.5 0.13
a.flags.writeable 8 107.4 18.9 0.18
np.require(a, ['A', 'W']) 8 5.21 4.58 0.88
np.broadcast_arrays(a, b) 8 1.20 1.11 0.92
np.unique(small) 8 2.60 2.41 0.93
np.linspace(0, 1, 5) 8 0.517 0.516 1.00
np.delete(a, 1) 8 2.28 2.27 1.00

2. With the fix (3.14t, main at f99b1a0 vs main + fix)

Workload Threads main Fix Fix / main
a.flags 1 19.5 18.4 0.94
a.flags 2 10.6 36.8 3.48
a.flags 4 15.3 73.2 4.77
a.flags 8 19.6 123.5 6.29
a.flags.writeable 8 18.9 92.9 4.91
np.require(a, ['A', 'W']) 8 5.22 5.39 1.03
np.broadcast_arrays(a, b) 8 2.83 2.79 0.99
np.unique(small) 8 2.49 2.58 1.04
np.linspace(0, 1, 5) 8 0.863 0.877 1.02
np.delete(a, 1) 8 2.17 2.25 1.04
a.shape (control) 8 162.3 144.2 0.89

3. Single-thread cost on a GIL build (3.14, ns per call, 5 runs per build)

Workload main Fix Difference
a.shape (control, not touched) 45.6 48.1 +2.6 ns
np.shape(a) (calls a dispatcher, creates none) 143.3 146.8 +3.5 ns
a.flags 43.9 49.2 +5.2 ns
a.flags.writeable 64.3 71.3 +7.0 ns
np.require(a, ['A', 'W']) 672 688 +2.3%
np.broadcast_arrays(a, b) 1325 1360 +2.7%

AI Disclosure

To implement the changes and benchmark with claude opus 5.5 max, I have reviewed the code.

@seberg

seberg commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

@kumaraditya303 is this a sane approach? I guess it is fine at least when subclassing isn't possible (maybe generally, but dunno), but it feels like a hack (i.e. it's just not the job of tp_free to do this...!?).

(Basically, this is probably fine, with a code comment, but it piles on that we are working around something missing in CPythong: Proper heap type immortalization and maybe even TypeDecref/Incref() API or so!? NumPy can't be the only project that has these issues.
For the dtypes I am somewhat more happy with hacks, because they are at least so central that they are more likely to affect real world scenarios, these ones I am less sure about.)

@kumaraditya303

Copy link
Copy Markdown
Contributor

s this a sane approach? I guess it is fine at least when subclassing isn't possible (maybe generally, but dunno), but it feels like a hack (i.e. it's just not the job of tp_free to do this...!?).

This seems fine to me in the case of type which are not subclass-able, I think there could be issues if type is base type and subclass is implemented in C extension with a custom tp_dealloc but pure python subclasses would work.

@seberg

seberg commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Well, the main question is if this is worthwhile to worry about (in NumPy, I am assuming that CPython will provide us with better solutions eventually).
unique(small) being 7% slowdown with 8 threads doesn't seem shocking to me but is a real thing (all the benchmarks there are micro-benchmarks. arr.flags seems very unlikely to be "hot").

And I am tempted to say: This seems not like a big enough deal to do something that is only safe because we disallow subclasses, which we do here (and I am very sure tp_free isn't inherited by Python subclasses and I am not sure about C-ones).

I guess one could also call the super().tp_dealloc() to be less hacky, but I am not sure if that might not be more heavy weight to call (at least if we have to fetch the slot via the stable API function)?

EDIT: Another way to say why it's a weird: tp_dealloc pairs with tp_alloc, tp_free with tp_new. This wirese tp_new cleanup to tp_free, though.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants