perf: keep calls out of the tail-call instruction handlers - #77
Merged
explodingcamera merged 4 commits intoSep 29, 2026
Merged
explodingcamera merged 4 commits into
explodingcamera merged 4 commits into
Conversation
A tail-called handler that contains a call saves and restores a frame record on every instruction it executes, even when the call sits on a path that validated code never takes. Most handlers had such a call: the `instruction_handler_mismatch` and `stack_underflow` panics, the bounds-check panics of value-stack and global indexing, or the conversion of the `Trap::ValueStackOverflow` a push returned into an `ExecError`. Both tail-call dispatchers now `become` their handler-mismatch and instruction-pointer panics instead of calling them. Value-stack and global accesses that validation rules out stop through `invariant_violated`: with `nightly-tail-calls` in release builds that is `core::intrinsics::abort`, a trap instruction in place, and otherwise the panic it was before. A push inside a function body no longer returns a `Result`: `enter_locals` reserves the function's whole operand stack or traps before the body runs, so a full stack there is one of those invariants. Validated modules cannot break them; archives skip validation and must already come from a trusted source. Numeric, local, global and constant handlers no longer save a frame (`i32.add` is 26 instructions on aarch64, down from 33). Memory, call and return handlers still call out for their traps and slow paths.
`exec_load_local`, `exec_store_local_local`, `exec_inc_memory_local`, `exec_fma_store` and `exec_mem_load_lane` were the memory helpers without `#[inline(always)]`, and LLVM kept them out of line: each of their handlers made a call and ran the helper's own prologue and epilogue on every load or store. They are now inlined like `exec_mem_load` and `exec_mem_store`.
Owner
|
Amazing, thanks again for the detailed work on improving performance! I'm also seeing a lot of gains on my end. I might have to put I could also replicate frameless calls on x86, though in slightly less cases. |
`core::process::abort_immediate` is the stabilization track of the abort intrinsic. It inlines into the handlers the same way, so every handler still traps in place and keeps no frame.
The handlers reach the mismatch panic only through the cold `handler_mismatch` functions they tail-call, so the shared `instruction_handler_mismatch` helper is gone.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
With the tail-call dispatch, a handler that contains a call saves and restores a frame record on every instruction it executes, even when the call is on a path that validated code never takes. On
nextall but three of the 615 handlers had such a call: theinstruction_handler_mismatchandstack_underflowpanics, the bounds-check panics of value-stack and global indexing, or the conversion of theTrap::ValueStackOverflowa push could return into anExecError.The handlers now
becomethe mismatch and instruction-pointer panics instead of calling them. Value-stack and global accesses that validation rules out stop throughinvariant_violated, which withnightly-tail-callsin release builds iscore::intrinsics::abort(an inline trap instruction) and otherwise the panic it was before. A push inside a function body no longer returns aResult:enter_localsreserves the function's whole operand stack or traps before the body runs (#59). Only modules that skip validation can break these invariants, and archives, which do, are already documented as trusted input.The second commit inlines the five memory helpers that were not
#[inline(always)], so a load or store through a local address no longer callsexec_load_localand similar.448 of the 615 handlers are now frameless, and
i32.addis 23 instructions instead of 31. Memory, call and return handlers still keep a frame for their traps and slow paths.Change in cycles per call against
next(a0ea681), on the efficiency cores of an iPhone 12 (A14), iPhone XS Max (A12) and iPhone SE (A13), median of five interleaved launches per build:On an M4's efficiency cores these two commits give −8.6% instructions and −5.6% cycles, and the loop dispatch (without
nightly-tail-calls) −3.5% instructions and −1.1% cycles. Some gory details:I elaborated on the memory-order flushes in a separate PR. The good news is that the extra inlining isn't busting the cache, and the call avoidance isn't unraveling some other CPU performance in a way I didn't predict. we'll keep slamming into this memory-order flush stall pileup until we deal with it explicitly, but it's still net positive on wall clock and benchmarks even on the cut-down efficiency cores :)