Repository navigation
Conversation
round() ran multiply, rint and true_divide on the float16 array. Each pass stores its result as float16, so ``x * 10**decimals`` became inf above 65504, and the scale factor itself is already inf in float16 for decimals >= 5, which gave nan/inf for values such as ``np.round(np.float16(2.0), 5)`` (numpygh-13699). Intermediate results were also rounded to float16 before ``rint``. Round exact float16 arrays in float32 and convert back once, using the existing float32 code. Subclasses and an ``out`` of another dtype or shape keep using the old path.
The test converts every float16 bit pattern to float32, including signalling NaNs. On some platforms (ARM64 CI) that conversion raises "invalid", which the test configuration turns into an error. Do the conversion inside the existing errstate block.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR summary
Fixes #13699.
np.roundon a float16 array ranmultiply,rintandtrue_divideas three float16 ufunc passes. Since float16 has no arithmetic of its own, each pass converted every element float16 -> float32, computed, and converted the result float32 -> float16 again, so an element went through three round trips per call. Each of those conversions stores the result as float16, sox * 10**decimalsbecomesinfabove 65504, and the scale factor10**decimalsis alreadyinfin float16 fordecimals >= 5.This PR fix float16 case: it rounds float16 arrays in float32, using the existing float32 code, and converts back once. This leads to
decimal >= 5coverage (no longer warn overflow) and a small speedup. Others are unchanged.Checks I ran locally:
TestMethodsintest_multiarray.py.np.roundis about 3x faster than onmain; float32 and float64 are unchanged.Benchmark script
Run it once per NumPy build, pinned to one core, and compare:
bench_round_pr.py:compare_round_pr.py:Speedup (
maintime / PR time) on my machine (AVX2/FMA/F16C, one pinned core, 2 interleaved runs, min of each):AI Disclosure
Claude Code was used.