gh-156443: Keep PyLong loop carries as twodigits in shifts and division - #157060
gh-156443: Keep PyLong loop carries as twodigits in shifts and division#157060XiaohongGong wants to merge 7 commits into
Conversation
|
Most changes to Python require a NEWS entry. Add one using the blurb_it web app or the blurb command-line tool. If this change has little impact on Python users, wait for a maintainer to apply the |
|
Hi, I'm an engineer from NVIDIA that has signed the CLA and PSF. What should I do to pass the CLA check? |
Be careful with which email you sign the CLA (see https://devguide.python.org/getting-started/pull-request-lifecycle/#why-do-i-need-to-sign-the-cla-again) |
There was a problem hiding this comment.
Please remove the implementation details, and focus on user-facing changes. E.g., "speed up X on X".
There was a problem hiding this comment.
I'v updated the NEWS.d. Could you please take another look about that? Thanks for your comment!
Thanks for the comment! I confirmed that my company (NVIDIA) has signed the CLA and my email and github are all correct. Any other checks/steps that should I do to pass the CLA? |
|
You'll have to sign (again) by clicking the button above, I'm afraid otherwise we can't do anything here. |
If I sign again by clicking the button above, it means that I will sign on behalf on the individual contributor, which may not be recommended? The contribution is made on behalf of my company. My understanding is that NVIDIA has an existing PSF Contributor Agreement. Could you please let me know whether any additional action is needed to associate this PR/GitHub account with the corporate agreement? Note that I'v clicked the sign button above and the CLA check has passed. But I'm aware that should be a mistake and I should not do that. |
gh-156443: Keep PyLong loop carries as twodigits in shifts and division
Several functions in
longobject.c(v_lshift,v_rshift, andx_divrem) narrowed a loop carry value todigitorsdigit, then widened it again on the next iteration.On AArch64, that 64-to-32-to-64 conversion inserts an extra
movon the loop-carried critical path. Keeping the carry at twodigits until the function returns drops thatmovand shortens the carry chain. Results are unchanged for valid limbs.Use
v_lshiftas an example, the loop on AArch64 previously contains a redundantmov w3, w3on the carry chain:Keeping the carry as twodigits removes that narrowing conversion. The loop is optimized to:
The instruction count on the loop carried chain is reduced from 3 to 2. We can observe ~20% performance improvement of the
pyperformance pidigits benchmark on an NVIDIA Grace CPU, while no material regressions observed on other platforms and benchmarks.
Tests cover divmod of saturated limbs with quotients near BASE (x_divrem's inner loop), and intra-digit shifts via float() and true division: full-limb values and powers of ten.
Fixes gh-156443.
Co-authored-by: Kyrylo Tkachov ktkachov@nvidia.com