Skip to content

perf: keep calls out of the tail-call instruction handlers - #77

Merged
explodingcamera merged 4 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/frameless-handlers
Sep 29, 2026
Merged

explodingcamera merged 4 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/frameless-handlers

Conversation

@matthargett

Copy link
Copy Markdown
Contributor

With the tail-call dispatch, a handler that contains a call saves and restores a frame record on every instruction it executes, even when the call is on a path that validated code never takes. On next all but three of the 615 handlers had such a call: the instruction_handler_mismatch and stack_underflow panics, the bounds-check panics of value-stack and global indexing, or the conversion of the Trap::ValueStackOverflow a push could return into an ExecError.

The handlers now become the mismatch and instruction-pointer panics instead of calling them. Value-stack and global accesses that validation rules out stop through invariant_violated, which with nightly-tail-calls in release builds is core::intrinsics::abort (an inline trap instruction) and otherwise the panic it was before. A push inside a function body no longer returns a Result: enter_locals reserves the function's whole operand stack or traps before the body runs (#59). Only modules that skip validation can break these invariants, and archives, which do, are already documented as trusted input.

The second commit inlines the five memory helpers that were not #[inline(always)], so a load or store through a local address no longer calls exec_load_local and similar.

448 of the 615 handlers are now frameless, and i32.add is 23 instructions instead of 31. Memory, call and return handlers still keep a frame for their traps and slow paths.

Change in cycles per call against next (a0ea681), on the efficiency cores of an iPhone 12 (A14), iPhone XS Max (A12) and iPhone SE (A13), median of five interleaved launches per build:

benchmark A14 A12 A13
xmrsplayer (1024-frame buffer) −5.5% −4.0% −12.3%
audio DSP (1000 frames × 512) −3.8% −4.3% −9.7%
graphql-validation (AS) −4.6% −5.5% −7.7%
multi-memory twin: one memory −9.2% −8.7% −11.1%
crc32 (64 KB) −5.5% −10.4% −7.7%
convolution 256×256 −3.5% −5.9% −6.1%
sieve (10000) −14.1% −8.6% −13.9%
bulk_memory (memory.copy/fill) −6.7% −8.3% −7.6%
matmul relaxed-simd FMA −6.0% −8.4% −9.7%
GC binary trees (~130K struct.new) −1.3% −2.5% −1.9%
fib(30) −8.7% −6.4% −7.6%
tail-call FSM (65536 return_call) −8.3% −6.5% −7.7%
call_indirect (200K) −2.9% −5.9% −3.9%
call_ref (200K) −4.3% −3.8% −5.0%
vtable_poly4 (200K) −4.3% −4.1% −5.3%
EH parser, exnref (4096 stmts, 25% throw) −2.8% −1.4% −3.2%
geomean, cycles −5.8% −6.0% −7.6%
geomean, instructions −7.2% −7.2% −7.2%
geomean, wall time −5.7% −6.0% −7.7%
rows faster (cycles) 16/16 16/16 16/16

On an M4's efficiency cores these two commits give −8.6% instructions and −5.6% cycles, and the loop dispatch (without nightly-tail-calls) −3.5% instructions and −1.1% cycles. Some gory details:

  • xmrsplayer cycles per call fall from 39.5M to 37.4M, graphql from 32.3M to 30.8M.
  • branch mispredicts stay flat, within ±5%.
  • the main one cost is memory-order flushes: 7.7k → 13.2k per xmrsplayer call, roughly 0.3–0.4% of cycles against the 5.5% gain.

I elaborated on the memory-order flushes in a separate PR. The good news is that the extra inlining isn't busting the cache, and the call avoidance isn't unraveling some other CPU performance in a way I didn't predict. we'll keep slamming into this memory-order flush stall pileup until we deal with it explicitly, but it's still net positive on wall clock and benchmarks even on the cut-down efficiency cores :)

A tail-called handler that contains a call saves and restores a frame
record on every instruction it executes, even when the call sits on a path
that validated code never takes. Most handlers had such a call: the
`instruction_handler_mismatch` and `stack_underflow` panics, the bounds-check
panics of value-stack and global indexing, or the conversion of the
`Trap::ValueStackOverflow` a push returned into an `ExecError`.

Both tail-call dispatchers now `become` their handler-mismatch and
instruction-pointer panics instead of calling them. Value-stack and global
accesses that validation rules out stop through `invariant_violated`: with
`nightly-tail-calls` in release builds that is `core::intrinsics::abort`, a
trap instruction in place, and otherwise the panic it was before. A push
inside a function body no longer returns a `Result`: `enter_locals`
reserves the function's whole operand stack or traps before the body runs,
so a full stack there is one of those invariants. Validated modules cannot
break them; archives skip validation and must already come from a trusted
source.

Numeric, local, global and constant handlers no longer save a frame
(`i32.add` is 26 instructions on aarch64, down from 33). Memory, call and
return handlers still call out for their traps and slow paths.
`exec_load_local`, `exec_store_local_local`, `exec_inc_memory_local`,
`exec_fma_store` and `exec_mem_load_lane` were the memory helpers without
`#[inline(always)]`, and LLVM kept them out of line: each of their handlers
made a call and ran the helper's own prologue and epilogue on every load or
store. They are now inlined like `exec_mem_load` and `exec_mem_store`.
Comment thread crates/tinywasm/src/interpreter/executor/dispatch_become.rs
Comment thread crates/tinywasm/src/macros.rs Outdated
@explodingcamera

explodingcamera commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Amazing, thanks again for the detailed work on improving performance! I'm also seeing a lot of gains on my end. I might have to put abort_immediate() behind a feature flag later though if tail calls stabilize first / panicking instead of abort on unvalidated code has any use cases.

I could also replicate frameless calls on x86, though in slightly less cases.

`core::process::abort_immediate` is the stabilization track of the abort
intrinsic. It inlines into the handlers the same way, so every handler
still traps in place and keeps no frame.
The handlers reach the mismatch panic only through the cold
`handler_mismatch` functions they tail-call, so the shared
`instruction_handler_mismatch` helper is gone.
@explodingcamera
explodingcamera merged commit 017780e into explodingcamera:next Sep 29, 2026
11 checks passed
@matthargett
matthargett deleted the perf/frameless-handlers branch October 2, 2026 04:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants