Repository navigation
JIT: Implement unique reference tracking in Tier 2 for reference count optimizations #143414
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Jan 4, 2026 @Fidget-Spinner Hi, Could you please review this issue and let me know your thoughts?
- addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Jan 4, 2026 @cocolato seems correct to me. The key thing about the optimization is that
op(_UNPACK_SEQUENCE_TWO_TUPLE, (seq -- val1, val0)) { assert(oparg == 2); PyObject *seq_o = PyStackRef_AsPyObjectBorrow(seq); assert(PyTuple_CheckExact(seq_o)); DEOPT_IF(PyTuple_GET_SIZE(seq_o) != 2); STAT_INC(UNPACK_SEQUENCE, hit); val0 = PyStackRef_FromPyObjectNew(PyTuple_GET_ITEM(seq_o, 0)); val1 = PyStackRef_FromPyObjectNew(PyTuple_GET_ITEM(seq_o, 1)); PyStackRef_CLOSE(seq); }becomes
op(_UNPACK_SEQUENCE_TWO_TUPLE_UNIQUE_STEAL, (seq -- val1, val0)) { assert(oparg == 2); PyObject *seq_o = PyStackRef_AsPyObjectBorrow(seq); assert(PyTuple_CheckExact(seq_o)); DEOPT_IF(PyTuple_GET_SIZE(seq_o) != 2); STAT_INC(UNPACK_SEQUENCE, hit); val0 = PyStackRef_FromPyObjectSteal(PyTuple_GET_ITEM(seq_o, 0)); PyTuple_SET_ITEM(seq_o, 0, NULL); val1 = PyStackRef_FromPyObjectSteal(PyTuple_GET_ITEM(seq_o, 1)); PyTuple_SET_ITEM(seq_o, 1, NULL); PyStackRef_CLOSE_NO_ESCAPE(seq); }Which will allow the op to have no reference counts operations at all, and also be non-escaping. So we can stack cache over it.
Reacted by Hai Zhu@cocolato now that I think about this, this is more complicated than I thought:
The problem is you have to invalidate all unique references on any escaping uop, like see (
_PyUop_Flags[opcode] & HAS_ESCAPES_FLAG). RETURN_VALUE is an escaping uop, so we cannot easily apply this optimization across RETURN_VALUE.However, I think there's one useful place you can still apply this optimization: object creation and initialization. See the CALL_ALLOC_AND_ENTER_INIT instruction for example.
I think because of how complicated this is, I'll take over, sorry! However, I need your help on making this optimization better and can parallelize some work here: we need to specialize on more forms of object creation. Can you please add a specialization for
CALL_SLOT_AND_ENTER_INIT?Basically the current code JITs:
class A: def __init__(self, a): self.a = a def foo(n): for i in range(1, n + 1): x = A(1) return 1 foo(4003)But once you add
__slots__toA, there's no more specialization, and the JIT cannot optimize and just fails. So we need a new specialization for calling__slots__. This turns up frequently in dataclasses and also the bm_float benchmark on pyperformance.Just by adding
CALL_SLOT_AND_ENTER_INIT, you should see a speedup on that benchmark. This isn't an easy task, and I think you're a very capable/strong contributor, so I'm entrusting you with this. Do you mind taking it up? You can see an example of how to do add a specialization here #143389Sorry for discouraging you from this optimization btw. I feel with the current state of the JIT, we can't use this (yet). If you manage to implement more specializations, that should make it possible though.
Reacted by Hai ZhuWe can still do the optimizations, but we may have to insert extra guards. However, that's blocked by #143421 as well.
Let me know which one you want to work on, and I can let you take it up!The latter got taken up by Donghee, so it will have to be the new specialization!add a specialization for
CALL_SLOT_AND_ENTER_INITThanks for the explanation! I'm pleased to take on the task of adding a specialization for
CALL_SLOT_AND_ENTER_INIT. I'm also very willing to help with other related development tasks in the future—Truly I need to start with simpler tasks to get more familiar with this area of optimization. Thanks for the trust!Reacted by Ken Jin@cocolato I'm surprised, you're right. Really sorry, my mistake. I have another issue to work on. Can you please work on making the JIT optimization state per-thread?
The main one is
JitOptContextis stack allocated https://github.com/python/cpython/blob/main/Python/optimizer_analysis.c#L342We need it to be pushed to
_PyThreadStateImplhttps://github.com/python/cpython/blob/main/Include/internal/pycore_tstate.h#L149So that we don't stack overflow on the JIT. This time I've verified that the problem indeed exists. Do you want to take this one up?
Ok, I will take some time to understand the issue and work on it!
Reacted by Ken JinPlease don't work on the parent issue (
Make the JIT optimizer buffer add to a new buffer, not in-place), instead it's the sub-issue, thanks!Reacted by Hai Zhu@Fidget-Spinner Hi, are there some other tasks we can do for this issue?
2 remaining items
@cocolato We need to first implement a new symbolic type for objects. See optimizer_symbols.c and the things that are done there. I can let you take that up if you want.
Alright, you can do this now, let me know if you face any issues. Thanks!
I suspect we want to start simple. So just start with a symbolic type for objects with
__slots__, then you just need to maintain an index to symbol mapping.Recall that we do not want allocations in the JIT optimizer if we can avoid it. So I think it's okay to just store it as a pair in an array and do a linear search for now.
Reacted by Hai ZhuI suspect we want to start simple. So just start with a symbolic type for objects with
__slots__, then you just need to maintain an index to symbol mapping.cpython/Python/optimizer_symbols.c
Line 938 in 31c81ab
static_assert(sizeof(JitOptSymbol) <= 3 * sizeof(uint64_t), "JitOptSymbol has grown");
Why do we have this limitation here? If we now need to extendJitOptSymbolto add theJitOptSlotsObjecttype, this might restrict the length of the slot array we're tracking.@cocolato you can increase the size of the assert.
@cocolato on second thought, let's not change the size of that assert. That assert affects the size of all symbols. A possible solution without increasing the size is to contain a single pointer that points to an arena (see for example the t_arena and s_arena in the abstract optimizer struct). The arena itself will then be shared across all object types, and contain the slots/attributes for objects.
Reacted by Hai Zhu@cocolato sorry, do you mind if I let @reidenong do unique reference tracking while you're doing object property tracking? I think object property tracking will gain us huge immediate benefits. However, it's also quite large and complex. So I think you'll be occupied with it for awhile. In the meantime, we can parallelize some of this work.
@cocolato sorry, do you mind if I let @reidenong do unique reference tracking while you're doing object property tracking? I think object property tracking will gain us huge immediate benefits. However, it's also quite large and complex. So I think you'll be occupied with it for awhile. In the meantime, we can parallelize some of this work.
Thank you for letting me know! I don't mind at all. I will focus on object property tracking(but in the coming days, I may also have little time for development), and I'm sure I can learn a lot from other people's PRs in the meantime.
Reacted by Ken JinReacted by Ken Jin@reidenong @Fidget-Spinner I used the implementation of unique reference tracking from the PR to add inplace float operations to the jit. Results look promising on microbenchmarks:
Expression Speedup total += a*b + c: 2.1x total += a + b : 1.5x total += a*b + c*d: 1.6xOn nbody I measured a 15% speedup. Once the PR is merged I will rebase my work and do some more benchmarking.
Update: PR is at #146307
Reacted by Hai Zhu, Ken Jin and Chris Eibl- added a commit that references this issue
on Mar 22, 2026 @eendebakpt yeah that inplace op using
refcnt==1essentially was around in the older days (3.11-3.13) in the interpreter but removed due to it being not FT safe. It would be great to have that back in the JIT! I've merged the unique reference tracking PR, so please feel free to open a PR for float ops. Thanks!- added a commit that references this issue
on Mar 23, 2026 I guess any further issues can be done as follow-up. Thanks Hai Zhu and Reiden!
Reacted by Hai Zhu
Feature or enhancement
Proposal:
Motivation
We should implement
unique reference trackingin Tier 2 to facilitate optimizations that reduce reference counting overhead. For example, when a tuple is known to be uniquely referenced, we can "steal" its element references during unpacking without performing any reference counting operations.For reference: discussion in #142952
Technical Approach
REF_IS_UNIQUEbit (bit 1) to theJitOptRefunion inpycore_optimizer.h(code reference).PyJitRef_MakeUnique()andPyJitRef_IsUnique()helper functions.PyJitRef_StripReferenceInfoandJIT_BITS_TO_PTR_MASKEDto support this unique reference bit.Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
Linked PRs