Skip to content

perf: keep the value stack's height in a register across tail-call handlers - #79

Open
matthargett wants to merge 5 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/stack-height-arg
Open

matthargett wants to merge 5 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/stack-height-arg

Conversation

@matthargett

@matthargett matthargett commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

The stack-height half of #78, which I separate into four commits that can I measured one by one, plus a fix:

  1. Pick the tail-call handlers' calling convention per target (see below). Only what the convention passes in registers stays out of memory between two handlers.
  2. Keep each value stack's height outside its Vec. A lane's height is its own field over a Vec that only grows. Slots above the height are stale and written before they are read. This lets the handlers write the height back without unsafe. A function's operand-stack reservation is written along with its locals only when it's at most 64 slots. A larger one stays capacity, and each slot is written the first time a push reaches it, so a function whose deep branch rarely runs keeps no pages for it: 64 stores of a module whose untaken branch declares a 30,000-deep i64 stack keep next's footprint (35 KB each on macOS, 16–19 KB on glibc) instead of 273 KB and 250 KB. Only builds with debug assertions check indices against the height, and a new CI job runs the tests optimized with them.
  3. Hold the value stack in the executor while it runs. The executor takes the store's value stack when it's created and gives it back when it's dropped and around host calls, also when a host function catches a panic from a nested call (tested in tests/host_nested_panic.rs). Every handler receives the executor as a &mut, so a stack access needs one fewer load, and the compiler knows that no store to a slot or to memory changes a height (verified this in disassembly and with low-level CPU counters).
  4. Pass the 32-bit value stack's height through the tail-call handlers. A handler no longer loads the height its predecessor stored. The Unbudgeted handlers take six integer arguments and the Bounded ones five.
  5. Restore the store when a panic unwinds out of a call. On next, a host function that catches a panic from a nested call keeps that call's depth counted and its frames and values on the stacks, so after as many recoveries as the call stack allows, every nested call fails with CallStackOverflow. A panic caught around a root call leaves the store marked as executing. Guards now undo both however the call ends, and a resumable execution that a panic unwound out of stays completed instead of rerunning its slice. tests/host_nested_panic.rs catches 400 panics through Wasm frames with a call stack of 8 and value stacks of 16.

Apple devices

Change against next, median of five interleaved launches per build. Energy is the kernel's estimate of the CPU energy each call used (task_power_info_v2). Commit 1 does not change the code on 64-bit arm64. These were measured before commit 5 and commit 2's reservation change; on the iPhone SE, the updated series changes cycles by −0.1% and energy per call by −1.2% against them.

iPhone 12 (A14), efficiency cores iPhone SE (A13), efficiency cores iPhone XS Max (A12), efficiency cores Apple TV 4K (A10X)
commits 1–3, cycles −3.4% −7.5% −5.0% −4.0%
commits 1–3, energy per call −3.0% −4.4% −5.5%
commit 4 on top of 3, cycles −1.9% +1.0% −0.9% −4.2%
commit 4 on top of 3, energy per call −1.0% +1.2% +1.4%
all, cycles −5.2% (15/15 faster) −6.5% (15/15 faster) −5.9% (15/15 faster) −8.0% (14/15 faster)
all, energy per call −3.9% −3.2% −4.1%

On the A10X (Apple TV 4K first gen) the kernel doesn't reports energy or which cores ran the benchmark, which makes sense since it plugs into a wall :D I have an A10X iPad Pro, but that device didn't get updated past iPadOS 17 so the performance profiling is limited in different ways. So far I've found that what's good for A12 effiency cores maps well to A10X's high-performance (Hurricane) cores in terms of wallclock times.

On Apple Watch SE (S8, arm64_32), 13 benchmarks (the watch leaves out audio DSP and the GC trees due to peak memory usage constraints, something to optimize for later):

cycles wall time energy per call
commit 1 (extern "C-unwind") −13.2% (13/13 faster) −14.5%
all four commits −19.7% (13/13 faster) −19.8% −7.5%

Between some launches the watch changed its performance state, which moves energy per cycle between about 46 and 73 pJ, too often to give commit 1's energy on its own. Profiling on the physical watch is extremely annoying, and very touchy (literally!), but I feel like I got enough on-device data for it to be defensible.

benchmark A14 cycles A14 energy A13 cycles A13 energy
audio DSP −11.8% −7.7% −7.4% −1.4%
crc32 (scalar) −8.5% −5.2% −11.3% −5.3%
convolution (scalar) −8.0% −4.6% −8.6% −3.0%
xmrsplayer −6.2% −4.7% −5.1% −2.4%
matmul relaxed-simd FMA −5.9% −3.5% −6.4% −1.6%
tail-call FSM −5.8% −4.2% −7.2% −4.0%
call_ref twin: call_indirect −5.2% −5.1% −7.2% −4.5%
bulk_memory (scalar) −5.0% −4.2% −7.7% −4.0%
call_indirect −4.1% −4.3% −6.9% −4.3%
vtable_poly4 −3.8% −3.4% −6.8% −5.2%
graphql-validation −3.8% −2.5% −4.1% −1.4%
sieve (scalar) −3.6% −3.0% −8.8% −3.3%
EH parser, exnref −3.5% −3.3% −3.0% −3.1%
fib(30) −1.4% −1.8% −5.6% −3.0%
GC binary trees −1.1% −1.0% −1.5% −1.0%

Commit 4 pays where the height's store-to-load round trip costs: on the iPhone 12 it removes 88–97% of the memory-order flushes on every benchmark profiled, and on audio DSP back-end stalls fall from 14.5% to 3.0% of pipeline slots. On the iPhone SE and XS Max it costs about 1% in energy per call, while commits 1–3 save 4–6% there.

The call-heavy rows are bound by the call path instead. On the iPhone 12's tail-call FSM, about a third of the samples are the Shared<WasmFunction> refcount updates on each call and return, which #74 removes (−16.6% cycles and −13.4% energy per call on its own there). These builds target the A10, which has no LSE atomics, so each update is a load/store-exclusive loop.

Built for the A12 (LSE atomics) and tuned for the A13 instead, each refcount update is a single ldadd, and the whole stack gains more. Change against next built for the A10:

iPhone 12 (A14) iPhone SE (A13) iPhone XS Max (A12)
cycles −14.9% (15/15 faster) −13.6% (15/15 faster) −9.0% (14/15 faster)
energy per call −11.0% −8.2% −7.6%

The A12 target alone makes next 4–11% faster, all of it on the call-heavy rows, and does nothing for #74, which removes those updates.

With commit 3 the stable dispatch runs 2–9% fewer instructions on every benchmark, 5–7% on most.

Calling convention (commit 1)

  • x64: the handlers are extern "rust-preserve-none", a nightly feature (rust_preserve_none_cc) enabled only with nightly-tail-calls, which already needs nightly for become. The Rust ABI passes six integer arguments in registers on SysV but four on Windows, so on Windows every dispatch puts the Instruction on the stack today, and with commit 4 the height too. extern "sysv64" didn't appear help: it passed the packed Instruction on the stack. rust-preserve-none passes twelve in registers on every OS and keeps only RSP and RBP callee-saved.
  • arm64_32 (watchOS 11): the handlers are extern "C-unwind". The Rust ABI passes the 8-byte Instruction by reference there, so on next every dispatch stores it and the next handler loads it back. The C convention passes it in a register, and C-unwind still lets a host function's panic unwind (super cool!)
  • Everything else keeps the Rust ABI.

Hot path per dispatch, weighted by xmrsplayer's dispatch mix, from the compiled handlers (instructions / stack operands / pushes and pops):

x86-64 System V Windows x64
next 28.0 / 0.43 / 1.13 31.5 / 1.88 / 2.99
commit 4 alone (Rust ABI) 29.0 / 0.38 / 1.24 34.0 / 3.49 / 3.29
all four commits 27.0 / 0.38 / 0.06 27.2 / 0.37 / 0.05

Handlers that push anything: 297 → 117 on System V and 556 → 54 on Windows. These are counts from the compiled handlers; they have not been timed on x64. I have a Surface Book 2 here I can resurrect and test on if need be, but I'd love to outsource to other people's deployment targets that I surely don't have :)

The Unbudgeted handlers take five integer arguments. The Rust ABI passes
six in registers on x86-64 System V but four on Windows, which puts the
Instruction on the stack on every dispatch there. `rust-preserve-none`
(nightly, like `become`) passes twelve in registers on every OS and keeps
only RSP and RBP callee-saved, so a handler saves at most RBP.

On arm64_32 (watchOS) the Rust ABI passes the 8-byte Instruction by
reference, so every dispatch stores it and the next handler loads it back.
The C convention passes it in a register, and `C-unwind` still lets a
host function's panic unwind.

Other targets keep the Rust ABI.
The height of a value-stack lane is its own field, over a Vec that holds
every slot the lane has reached. Slots above the height hold stale values
that are written before they are read again.

The tail-call handlers can then pass the height between them in a register
and write it back without unsafe code.

Entering a function writes its locals, and its operand-stack reservation
if that is at most 64 slots, so later entries at that height compare once.
A larger reservation stays capacity, and a push writes each slot the first
time it reaches it (on the tail-call build with push_within_capacity, which
makes no call), so a function whose deep branch rarely runs keeps no pages
for it. With 64 stores of a module whose untaken branch declares a
30000-deep i64 stack, each store keeps next's footprint (35 KB on macOS,
16-19 KB on glibc) instead of 273 KB and 250 KB.

v128.any_true and i8x16.all_true test the vector with integer arithmetic:
LLVM scalarized their iterator forms once pushes gained that path.

Only builds with debug assertions check indices against the height. CI
also runs the tests optimized with them, keeping the release profile's
wrapping arithmetic.

Entering a function zeroes one or two locals with direct stores, and only
the ones in between with memset.
The executor takes the store's value stack when it is created and gives it
back when it is dropped, and around host calls, which read their arguments
from the store and may call back into wasm.

Every handler receives the executor as a `&mut`, so a stack access reached
through it needs one load fewer (the store pointer), and the compiler knows
that no store to a slot or to memory changes a stack's height, so it can
keep the heights in registers within a handler.
…lers

Each tail-call handler receives the height of the 32-bit value stack, which
most instructions push to or pop from, as its last argument. It writes the
height to the stack before its body and reads it back before dispatching,
and with the stack in the executor the compiler can keep it in a register
in between, as it does in most handlers. A handler no longer loads the
height its predecessor stored, a store-to-load round trip through memory
on every dispatch.

The Unbudgeted handlers now take six integer arguments and the Bounded ones
five. With the calling conventions from the first commit they all stay in
registers: arm64 passes eight, and `rust-preserve-none` twelve on x86-64.
A host function that catches a panic from a nested call went on with that
call's depth still counted and its frames and values still on the stacks.
After as many recoveries as the call stack allows, every nested call failed
with CallStackOverflow. A panic caught around a root call left the store
marked as executing, so it rejected every later call.

Guards now undo both whether the call returns or unwinds. A resumable
execution that a panic unwound out of stays completed: its frame no longer
matches the stacks, so resuming it returns an error instead of running the
slice again. Dropping it clears the stacks, which would otherwise keep the
unwound call's values as GC roots until the store's next call.

The tests catch repeated panics through Wasm frames with a call stack of 8
and value stacks of 16, call a store again after catching panics out of
root calls, and resume after a caught panic.
@explodingcamera
explodingcamera self-requested a review October 1, 2026 16:09
explodingcamera added a commit that referenced this pull request Oct 3, 2026
originally reported in #79

Signed-off-by: Henry <mail@henrygressmann.de>
explodingcamera added a commit that referenced this pull request Oct 3, 2026
based on the fix in #79

Signed-off-by: Henry <mail@henrygressmann.de>
@explodingcamera

explodingcamera commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

I just pushed a fix for the function call panics from this to next. I do see a good performance improvement overall, and I’d definitely like to bring in more of these changes. I’m just a bit hesitant to add more nightly features right now, so I might shuffle the flags around to separate tail-call dispatch from a more general nightly feature.

I’ll probably do some more testing in this direction and pull the changes in incrementally rather than all at once.

@explodingcamera
explodingcamera removed their request for review October 3, 2026 14:32

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants