perf: keep the value stack's height in a register across tail-call handlers - #79
Open
matthargett wants to merge 5 commits into
Open
matthargett wants to merge 5 commits into
matthargett wants to merge 5 commits into
Conversation
The Unbudgeted handlers take five integer arguments. The Rust ABI passes six in registers on x86-64 System V but four on Windows, which puts the Instruction on the stack on every dispatch there. `rust-preserve-none` (nightly, like `become`) passes twelve in registers on every OS and keeps only RSP and RBP callee-saved, so a handler saves at most RBP. On arm64_32 (watchOS) the Rust ABI passes the 8-byte Instruction by reference, so every dispatch stores it and the next handler loads it back. The C convention passes it in a register, and `C-unwind` still lets a host function's panic unwind. Other targets keep the Rust ABI.
The height of a value-stack lane is its own field, over a Vec that holds every slot the lane has reached. Slots above the height hold stale values that are written before they are read again. The tail-call handlers can then pass the height between them in a register and write it back without unsafe code. Entering a function writes its locals, and its operand-stack reservation if that is at most 64 slots, so later entries at that height compare once. A larger reservation stays capacity, and a push writes each slot the first time it reaches it (on the tail-call build with push_within_capacity, which makes no call), so a function whose deep branch rarely runs keeps no pages for it. With 64 stores of a module whose untaken branch declares a 30000-deep i64 stack, each store keeps next's footprint (35 KB on macOS, 16-19 KB on glibc) instead of 273 KB and 250 KB. v128.any_true and i8x16.all_true test the vector with integer arithmetic: LLVM scalarized their iterator forms once pushes gained that path. Only builds with debug assertions check indices against the height. CI also runs the tests optimized with them, keeping the release profile's wrapping arithmetic. Entering a function zeroes one or two locals with direct stores, and only the ones in between with memset.
The executor takes the store's value stack when it is created and gives it back when it is dropped, and around host calls, which read their arguments from the store and may call back into wasm. Every handler receives the executor as a `&mut`, so a stack access reached through it needs one load fewer (the store pointer), and the compiler knows that no store to a slot or to memory changes a stack's height, so it can keep the heights in registers within a handler.
…lers Each tail-call handler receives the height of the 32-bit value stack, which most instructions push to or pop from, as its last argument. It writes the height to the stack before its body and reads it back before dispatching, and with the stack in the executor the compiler can keep it in a register in between, as it does in most handlers. A handler no longer loads the height its predecessor stored, a store-to-load round trip through memory on every dispatch. The Unbudgeted handlers now take six integer arguments and the Bounded ones five. With the calling conventions from the first commit they all stay in registers: arm64 passes eight, and `rust-preserve-none` twelve on x86-64.
A host function that catches a panic from a nested call went on with that call's depth still counted and its frames and values still on the stacks. After as many recoveries as the call stack allows, every nested call failed with CallStackOverflow. A panic caught around a root call left the store marked as executing, so it rejected every later call. Guards now undo both whether the call returns or unwinds. A resumable execution that a panic unwound out of stays completed: its frame no longer matches the stacks, so resuming it returns an error instead of running the slice again. Dropping it clears the stacks, which would otherwise keep the unwound call's values as GC roots until the store's next call. The tests catch repeated panics through Wasm frames with a call stack of 8 and value stacks of 16, call a store again after catching panics out of root calls, and resume after a caught panic.
matthargett
force-pushed
the
perf/stack-height-arg
branch
from
September 30, 2026 21:32
7bb7a50 to
29679a7
Compare
explodingcamera
self-requested a review
October 1, 2026 16:09
explodingcamera
added a commit
that referenced
this pull request
Oct 3, 2026
originally reported in #79 Signed-off-by: Henry <mail@henrygressmann.de>
explodingcamera
added a commit
that referenced
this pull request
Oct 3, 2026
based on the fix in #79 Signed-off-by: Henry <mail@henrygressmann.de>
Owner
|
I just pushed a fix for the function call panics from this to next. I do see a good performance improvement overall, and I’d definitely like to bring in more of these changes. I’m just a bit hesitant to add more nightly features right now, so I might shuffle the flags around to separate tail-call dispatch from a more general nightly feature. I’ll probably do some more testing in this direction and pull the changes in incrementally rather than all at once. |
explodingcamera
removed their request for review
October 3, 2026 14:32
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The stack-height half of #78, which I separate into four commits that can I measured one by one, plus a fix:
Vec. A lane's height is its own field over aVecthat only grows. Slots above the height are stale and written before they are read. This lets the handlers write the height back withoutunsafe. A function's operand-stack reservation is written along with its locals only when it's at most 64 slots. A larger one stays capacity, and each slot is written the first time a push reaches it, so a function whose deep branch rarely runs keeps no pages for it: 64 stores of a module whose untaken branch declares a 30,000-deep i64 stack keepnext's footprint (35 KB each on macOS, 16–19 KB on glibc) instead of 273 KB and 250 KB. Only builds with debug assertions check indices against the height, and a new CI job runs the tests optimized with them.tests/host_nested_panic.rs). Every handler receives the executor as a&mut, so a stack access needs one fewer load, and the compiler knows that no store to a slot or to memory changes a height (verified this in disassembly and with low-level CPU counters).next, a host function that catches a panic from a nested call keeps that call's depth counted and its frames and values on the stacks, so after as many recoveries as the call stack allows, every nested call fails withCallStackOverflow. A panic caught around a root call leaves the store marked as executing. Guards now undo both however the call ends, and a resumable execution that a panic unwound out of stays completed instead of rerunning its slice.tests/host_nested_panic.rscatches 400 panics through Wasm frames with a call stack of 8 and value stacks of 16.Apple devices
Change against
next, median of five interleaved launches per build. Energy is the kernel's estimate of the CPU energy each call used (task_power_info_v2). Commit 1 does not change the code on 64-bit arm64. These were measured before commit 5 and commit 2's reservation change; on the iPhone SE, the updated series changes cycles by −0.1% and energy per call by −1.2% against them.On the A10X (Apple TV 4K first gen) the kernel doesn't reports energy or which cores ran the benchmark, which makes sense since it plugs into a wall :D I have an A10X iPad Pro, but that device didn't get updated past iPadOS 17 so the performance profiling is limited in different ways. So far I've found that what's good for A12 effiency cores maps well to A10X's high-performance (Hurricane) cores in terms of wallclock times.
On Apple Watch SE (S8, arm64_32), 13 benchmarks (the watch leaves out audio DSP and the GC trees due to peak memory usage constraints, something to optimize for later):
extern "C-unwind")Between some launches the watch changed its performance state, which moves energy per cycle between about 46 and 73 pJ, too often to give commit 1's energy on its own. Profiling on the physical watch is extremely annoying, and very touchy (literally!), but I feel like I got enough on-device data for it to be defensible.
Commit 4 pays where the height's store-to-load round trip costs: on the iPhone 12 it removes 88–97% of the memory-order flushes on every benchmark profiled, and on audio DSP back-end stalls fall from 14.5% to 3.0% of pipeline slots. On the iPhone SE and XS Max it costs about 1% in energy per call, while commits 1–3 save 4–6% there.
The call-heavy rows are bound by the call path instead. On the iPhone 12's tail-call FSM, about a third of the samples are the
Shared<WasmFunction>refcount updates on each call and return, which #74 removes (−16.6% cycles and −13.4% energy per call on its own there). These builds target the A10, which has no LSE atomics, so each update is a load/store-exclusive loop.Built for the A12 (LSE atomics) and tuned for the A13 instead, each refcount update is a single
ldadd, and the whole stack gains more. Change againstnextbuilt for the A10:The A12 target alone makes
next4–11% faster, all of it on the call-heavy rows, and does nothing for #74, which removes those updates.With commit 3 the stable dispatch runs 2–9% fewer instructions on every benchmark, 5–7% on most.
Calling convention (commit 1)
extern "rust-preserve-none", a nightly feature (rust_preserve_none_cc) enabled only withnightly-tail-calls, which already needs nightly forbecome. The Rust ABI passes six integer arguments in registers on SysV but four on Windows, so on Windows every dispatch puts theInstructionon the stack today, and with commit 4 the height too.extern "sysv64"didn't appear help: it passed the packedInstructionon the stack.rust-preserve-nonepasses twelve in registers on every OS and keeps only RSP and RBP callee-saved.extern "C-unwind". The Rust ABI passes the 8-byteInstructionby reference there, so onnextevery dispatch stores it and the next handler loads it back. The C convention passes it in a register, andC-unwindstill lets a host function's panic unwind (super cool!)Hot path per dispatch, weighted by xmrsplayer's dispatch mix, from the compiled handlers (instructions / stack operands / pushes and pops):
nextHandlers that push anything: 297 → 117 on System V and 556 → 54 on Windows. These are counts from the compiled handlers; they have not been timed on x64. I have a Surface Book 2 here I can resurrect and test on if need be, but I'd love to outsource to other people's deployment targets that I surely don't have :)