RFC: pass a stack pointer through handlers to avoid loads and round trips #78
Replies: 2 comments
|
Yeah, this is interesting. I’d like to try passing the stack pointer through the handlers, though I’m a little worried about adding another argument when I’d also like to keep room in registers for accumulators later. No need to work around the acc branch, that was (outside of improvements to the parser that i pulled from there) a failed experiment 😄. My guess is that tinywasm’s fused opcodes already skip a lot of stack accesses, and my changes got in the way of that. I've tried caching the top of the stack in the past too, but the extra branch in the common path wiped out the benefit. I think I’ll need to rethink the parser and IR more thoroughly to make accumulators work, so it’d be useful to measure the stack-pointer idea separately. |
|
awesome. I measured both concerns you ran into during your experiments: the registers and the branch. Register budgetA probe with 24 integer and 24 float arguments per function, compiled for each target, counts how many arrive in registers:
The
Top of stackI built a reduction example in Rust with eight handler kinds, up to 64 copies each, dispatched like tinywasm. Only where the stack's height and top live varies. I ran it on an M4's efficiency cores and on the efficiency cores of an iPhone 12, SE and XS Max. Cycles per dispatch at 171 handlers (the real loops I profile use 136–226 handler pairs):
In the hot loops I profile, nearly every dispatch is a 32-bit lane op or control flow. So the two registers to add are that lane's height and top: 7 of 8 on arm64. On x64 they fit with In other words, the stack's height and top get home registers, the way wasm3's ops all receive The model's handlers are simpler than tinywasm's, and the fused ops already skip the stack a lot, so the real-world gain in integrated will be well below these numbers. I can prototype the height first, so it's measured on its own as you suggested, and then add the top. Just lmk how you'd like me to proceed based on this data! |
Uh oh!
There was an error while loading. Please reload this page.
I made this an RFC because it looks like the
accbranch is already poking around in this area. I wanted to get some exhaustive numbers up to help drive those decisions.Problem
the flushes themselves are only 1.8% of pipeline slots in the model and 2.4% in tinywasm. Most of the cost is loads waiting on older stores. At 64 copies, back-end stalls rise from 5.3% to 25.3% of slots, and execution latency is 17× higher. In tinywasm they rise from 13.2% to 19.7% with 3× the execution latency. That matches iPhone 12, where audio DSP gained 89M back-end cycles per call with our changes.
Proposal
passing a stack pointer through the handlers, the way #75 passes the instruction slice. this should make every stack address computable without a load, as well as leaving nothing for the predictor to guess. A top-of-stack register should also remove the top value's round trip. I can submit a PR for this, but I didn't want to stomp on the
accbranch activity.Evidence
To prove/disprove, I made a reduction model of tinywasm's dispatch: K copies of one handler that updates the stack top in place, each copy running the same 25 instructions and dispatching with
becomethrough a table, in a loop.with no aliasing between handlers, 64 copies run at 5.28 cycles per dispatch and never flush. In this testing, I figured out that live pairs count, not total pairs! the same 64 copies run in blocks of 64 (few pairs active at once) flush 0.12 times per 1,000 dispatches.
Inside tinywasm is a loop whose statements cycle through
Ddistinct stack binops, so 5 to 45 distinct handler pairs, with instructions per dispatch unchanged. Flushes and cycles are flat up to 11 pairs, step up between 11 and 19, then plateau. Our faster frameless handlers from #77 pay more: +9.4% cycles across the range, against +5.2% on #75. The real loops in my benchmark suite are well past that: Audio DSP needs 40 distinct handler pairs for 90% of its dispatches, xmrsplayer 136 and graphql 226.so this refines the original guess: the flushes themselves are only 1.8% of pipeline slots in the reduction and 2.4% in tinywasm. Most of the cost is loads waiting on older stores. At 64 copies, back-end stalls rise from 5.3% to 25.3% of slots, and execution latency is 17× higher. In tinywasm they rise from 13.2% to 19.7% with 3× the execution latency. That M4 e-core result matches the iPhone 12, where audio DSP gained 89M back-end cycles per call with our changes.
All reactions