Skip to content

Worker tax 1.47x -> 1.04x (Linux), 1.91x -> 1.07x (macOS): per-thread allocator, TLS blocks, one exception alert - #812

Merged
ctate merged 10 commits into
mainfrom
perf/workers/01-worker-tax
Oct 10, 2026
Merged

ctate merged 10 commits into
mainfrom
perf/workers/01-worker-tax

Conversation

@cramforce

Copy link
Copy Markdown
Contributor

A program that merely contains new Worker(...) ran its single-threaded code much slower than the same program without one: with origin/main's compiler the tsc-ts check took 1.91x as long on macOS and 1.47x on Linux when built as a Worker program. With this branch:

Changes (worker executables only unless noted):

  • A small-object allocator per script thread; the emitted allocation fast path reads one thread-local block.
  • int32 class-field specialization also in worker executables.
  • ELF worker executables use local-exec/initial-exec TLS, and ELF worker runtime units compile with -ftls-model=initial-exec.
  • Emitted exception checks test one process-wide alert word instead of the thread-local active cell and stop signal. A thread raises its share when its active cell may become pending (throws, rethrows, the termination sentinel, fiber switches, moves into the cell) and drops it once the cell is clear; worker.terminate() holds a reference until the worker's script ends.
  • Termination polling treats more operations as bounded (Node-style throws, runtime fences, array truncation, typed-array stores, static Map/Set/typed-array construction, caught-value and TDZ tests, primitive-argument string.*/num.* calls).
  • Thread-owned typed arrays are read and written inline.
  • The allocator's and the cycle collector's per-thread state each live in one thread-local block.

Tests: corpus native-worker-typed-arrays.ts; may-throw unit test for the bounded kinds; worker-threads codegen tests. Runtime benchmark suite on Linux vs origin/main (4 layouts × 12 runs): all 17 workloads neutral, size +0.0%, geomean +0.9%.

Numbers: tsc-ts checking itself. Linux: exclusive 8-vCPU sandbox; macOS: shared, loaded machine (earlier quieter runs given where available).

Worker builds skipped the whole-program int32 field analysis, as library
builds do, so a Worker-using tsc-ts kept every flags field a double and
isSimpleTypeRelatedTo tripled in size. Unlike library instances, worker
instances never cross a native boundary: they reach other threads only
through postMessage and workerData, whose dynamic views read and commit
fields like every other dynamic reader the analysis already handles.

worker.new and the ArrayBuffer/SharedArrayBuffer constructors take
dynamic arguments but never store into a live class capsule (none of
their runtime paths reaches scr_dyn_typed_ref_commit), so they join the
read-only helpers instead of disqualifying every field a dynamic store
could name.
Worker executables compiled the small-object allocator out and sent
every allocation and free through the system allocator, because
several script threads would race on one set of free lists. Each script
thread now owns an arena: scr_sa is thread-local in worker builds and
describes one of 32 equal slices of a single process-wide reservation,
so the owning arena of any block is address arithmetic. A thread claims
an arena on its first allocation and parks it, free lists and bump
pointers intact, when it exits; the next thread continues from there.
Memory is never unmapped, so blocks a worker leaves behind (published
immortal objects) stay valid after it exits. A block freed by another
thread goes to its owner's lock-free remote list, drained on the owner's
slow path; realloc moves foreign blocks into the caller's arena.

The emitted constructor and release fast paths are inlined for worker
programs too, addressing the thread-local allocator state, live count
and dispose hook through llvm.threadlocal.address.

The allocator unit test gains a worker build that hands blocks across
threads and reclaims a parked arena, also under ThreadSanitizer.
Every thread-local of a worker executable lives in the executable's own
TLS block: the program's internal ones and the runtime's external ones
alike. The emitter declared them with LLVM's default general-dynamic
model, which the linker relaxes to `mov %fs:0` plus an add at each
access. On ELF targets the program's thread-locals are now local-exec
and the runtime's initial-exec, so each access is one %fs-relative
address that LLVM folds into the load or store. Darwin (one model),
Windows, WASI and shared-library thread instances are unchanged.

Linux tsc-ts in a Worker build: 5.01 s -> 4.96 s (7 interleaved runs,
non-worker build 4.55 s); the 16.5k relaxed sequences became direct
%fs accesses.
Every emitted pending check of a worker executable loaded the
thread-local active exception cell and the thread-local stop signal.
On Darwin each thread-local access is a call (_tlv_get_addr), and LLVM
does not share them across a function's checks: compareTypes made 61
such calls, and the single-threaded tsc-ts self-check spent 19% of its
time in them (a Worker build ran 1.67x slower than the same program
without Worker on macOS).

The checks now load one ordinary global, scr_exc_alert, which is
nonzero while any script thread may have an exception pending in its
active cell or must stop. Zero proves nothing is pending anywhere, so
the check falls through; otherwise the out-of-line scr_exc_pending
answers for the calling thread, observes termination and settles the
thread's share. Each thread holds at most one share: raised by throws,
rethrows, the termination sentinel, fiber switches and payload moves
into the active cell, dropped by take, clear, moves out of the active
cell and the slow path. A share is only changed by its own thread, so a
thread's raise is visible to all of its later checks. worker.terminate()
holds one more reference until the worker's script has ended, so the
worker reaches its slow path and stops. An exception in flight in one
thread costs the others slow-path calls only until it is caught.

The stop signal is runtime-private again; the active cell stays
exported for ordinary executables' inline checks.

tsc-ts self-check, macOS, Worker build vs non-worker build: 1.67x ->
1.18x (check phase 1.65x -> 1.09x).
A worker function polls for termination on entry (and so becomes
may-throw, with checks after every call to it) unless all of its
operations are on the bounded list. The list missed operations that run
no script code and cannot block: compiler-resolved Node errors
(error.nodeThrow, in 3,986 tsc-ts functions), runtime fences, array
truncation, typed-array stores and reads, Map and Set construction over
literal entries or static containers, typed-array construction from
static sources, caught-value tests, module TDZ checks, and string and
number library calls whose arguments are primitives (a dynamic argument
could reach a user conversion).

tsc-ts: entry polls 6,903 -> 4,447 of 11,464 functions; the remainder
are almost all members of call cycles.
Worker programs sent every typed-array element read and write through
scr_bytes_get/scr_bytes_set, which take the shared-memory lock guard
even for buffers no other thread can see (4% of a Worker build of the
tsc-ts self-check). A view's `shared` storage is fixed before script
code can see it, so the emitted access now tests it: SharedArrayBuffer
views keep the locked runtime call, every other view takes the ordinary
inline path of non-worker programs.

The new corpus program reads and writes own and shared views through
the same functions (clamping, wrapping, float narrowing, out-of-range
and negative indices) on both threads.
In worker executables the small-object allocator state, the collector's
live and old-freed counts and the weak dispose hook were four separate
thread-locals, and every allocation or free touched them all. On
Darwin each distinct thread-local costs a _tlv_get_addr call per
function, so a runtime free made four such calls.

They now share one ScrThreadHot block (macros keep the runtime sources
unchanged; ordinary and library builds keep separate globals). The
emitted inline paths address the fields at fixed offsets from a single
llvm.threadlocal.address, which LLVM shares across a function.

tsc-ts self-check, macOS, Worker build vs non-worker build: 1.12x ->
1.10x.
The collector's thread-locals in worker executables (disposal queue,
candidate buffers, pass vectors, generation limit and scheduling state)
were twenty separate variables. scr_rc_destroy touched four of them, and
the release and walk paths several more; on Darwin each costs a
_tlv_get_addr call per function. Worker builds now keep them in one
ScrCycThread block (macros keep the sources unchanged); ordinary and
library builds keep separate variables.

tsc-ts self-check, macOS, Worker build vs non-worker build: 1.10x ->
1.08x.
Worker runtime variants link only into executables, whose own TLS block
holds every runtime thread-local, but they were compiled with the
default general-dynamic model. The linker relaxes each __tls_get_addr
call to a %fs-relative address, yet the compiler had already treated it
as a call: hot runtime leaves saved and restored extra registers around
every thread-local access (scr_mg_visit pushed five registers instead
of one). With -ftls-model=initial-exec each access is one %fs-relative
operand. Darwin, Windows and wasm variants are unchanged.

tsc-ts self-check, Linux sandbox, Worker build vs non-worker build:
1.04x -> 1.03x (9 interleaved runs).
@vercel

vercel Bot commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
scriptc Ready Ready Preview, v0 Oct 10, 2026 5:56am UTC

@ctate
ctate merged commit 33b040d into main Oct 10, 2026
44 of 45 checks passed

This branch was successfully deployed

1 active deployment
Preview — 1767fe7f Deployed Oct 10, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants