perf(qwp): improved Gorilla timestamp decoding - #69
glasstiger wants to merge 4 commits into
Conversation
decode() and decodeView() each held a copy of a Gorilla loop that read the bitstream one bit at a time and built every delta-of-delta from per-bit BigInt operations. Both now call one decoder that reads through a 32-bit bit buffer, carries the int64 delta and timestamp as int32 halves with explicit carries, wrapping exactly as BigInt.asIntN(64, ...) does, and emits runs of repeated deltas in one step. Values are stored as int32 word pairs through an Int32Array, with a byte swap on big-endian hosts so the bytes stay little-endian. Decoded values and errors are unchanged. A stream too short to hold one bit per row is now rejected before any storage is sized for it, and decode() reuses one decode buffer per decoder across batches, as column views already did. On 10,000-row batches in Node 24, column views decode 3.2-32x faster and materialized batches 1.8-27x faster, least for a constant interval and most for irregular timestamps. JavaScriptCore views decode 66-107x faster. Adds code-width, NULL, int64-edge, and decoder-reuse tests for both paths, a 600-stream comparison against a bit-at-a-time BigInt reference decoder, every malformed-column error, and timestamp-only Gorilla arms in benchmarks/egress.bench.ts.
The code-width test reused one decoder only with shorter batches, so nothing exercised growing the retained Gorilla storage: a decoder that stopped growing it passed every test while a larger later batch threw a RangeError. A three-value batch now precedes the main frame, which has to grow that storage in both decode() and decodeView(). The test's `max` column claimed to cover the top of the int64 range but peaked about 7.1e11 below INT64_MAX. It now ends exactly at INT64_MAX.
The Gorilla decoder writes int32 words through an Int32Array and byte-swaps them on big-endian hosts, so column views still expose little-endian int64 bytes. No test reached that branch: NATIVE_LITTLE_ENDIAN is fixed when result-batch.ts loads and every CI host is little-endian, so deleting the swap call left all 897 QWP tests passing. The new test makes the module's Uint16Array.of(1) probe see big-endian storage, loads a fresh copy, and decodes through decode() and decodeView(). On a little-endian host the swap then shows as byte-reversed int32 halves in every value. It fails if the swap call is deleted or inverted, or if the swap drops a byte, shifts one wrong, or covers only half the words.
The decoder's comment said the speedup was least for a constant interval and most for irregular timestamps, in V8 and JavaScriptCore alike. In JavaScriptCore the jittered series gains least: two Bun 1.0.14 runs measured 66x and 61.7x for it, against 72x and 72.5x for a constant interval. The ordering now applies to V8 only, and the JavaScriptCore figure reads "over 60x" rather than a 65-110x range that the second run fell below.
|
Level 3 review — approve (head No admitted findings: Critical 0, Moderate 0, Minor 0; in-diff 0, out-of-diff breakage 0. Test gate passes with 0 admitted coverage gaps. The full suite passed (1,127 tests), along with package-boundary and browser suites, typechecks, lint, and formatting checks. Base/head benchmarks reproduced the claimed Gorilla decoding speedup. The decoder retains row-cap-bounded scratch storage; no regression was observed. Submodules: none. |
|
Level 3 review — approve (head No findings in this PR: Critical 0, Moderate 0, Minor 0. The test gate passes with no coverage gaps worth reporting. Evidence
Tradeoff: Pre-existing bug (not from this PR, not blocking; worth its own issue)
const decoder = new QwpResultBatchDecoder();
decoder.decodeView(message, 1); // slot 0 never used
decoder.resetQuerySchema(); // TypeError: Cannot read properties of undefined (reading 'release')The schema and batch-sequence reset after the loop then never run. Reproduced on base and head (Node 24, Bun 1.0.14). Suggested fix: Submodules: none. |
Summary
Gorilla-encoded
TIMESTAMP,TIMESTAMP_NANOS, andDATEcolumns in query results now decode without BigInt arithmetic, inpackages/client-core/src/_qwp/_core/result-batch.ts. In testing, a QuestDB nightly server flagged every result batch as Gorilla-encoded, so this path decodes the timestamp columns of typical query results in both@questdb/nodejs-clientand@questdb/browser-client. No public API change, and decoded values and errors are unchanged.decode()anddecodeView()each held a copy of a loop that read the bitstream one bit at a time and built every delta-of-delta from per-bit BigInt operations (value |= 1n << BigInt(bit)). Both now calldecodeGorillaTimestamps(), which reads through a 32-bit bit buffer, carries the int64 delta and timestamp as int32 halves with explicit carries (wrapping exactly asBigInt.asIntN(64, ...)does), and emits runs of repeated deltas in one step.Int32Array, which ran the decode loop 1.8-4.7x faster thanDataViewstores in V8. A byte-swap pass runs only on big-endian hosts.truncated QWP Gorilla bitstream).decode(). One buffer per decoder is reused across batches, as column views already did. Its BigInt read-out loop is deliberately separate fromreadInt64Values(): with one loop serving both, V8 ran this path up to 2x slower, depending on which caller it optimized first.decode()anddecodeView(); 600 random streams (including truncated ones and int64 wraparound) against a bit-at-a-time BigInt reference decoder, comparing values, consumed bytes, and errors; every malformed-column error through both paths.benchmarks/egress.bench.ts(constant, jittered, and irregular intervals).benchmarks/workloads.tsexports its existing xorshift generator for them.Performance
Diagnostic, not CI gates. Before is
main; two interleaved runs each, averaged.Node 24,
benchmarks/egress.bench.ts, one GorillaTIMESTAMPcolumn of 10k rows, batches/s:mainmainThe existing mixed batch (INT, DOUBLE, VARCHAR, constant-interval Gorilla TIMESTAMP): reusable column views ~13.5k → ~35.7k batches/s (2.6x),
decode()materialized 861 → 897 (1.04x); the traversal and Zstd arms do not exercise this code and are unchanged within noise.JavaScriptCore (Bun 1.0.14, Safari's engine), the built browser bundle, the same three interval shapes, 10k rows, median µs per batch:
mainmainThe gain is larger in JavaScriptCore because it allocates on every BigInt operation, while V8 lowers
BigInt.asIntN(64, ...)arithmetic to native int64 math. Materializing abigint[]still costs one BigInt allocation per value, which bounds the materialized speedup.Validation
pnpm exec vitest run(1,126 tests, including container integration),pnpm test:dist,pnpm test:qwp-browserpnpm typecheck,pnpm typecheck:test,pnpm typecheck:qwp-browser,pnpm typecheck:bench,pnpm typecheck:distpnpm eslint,pnpm lint:bench,pnpm format:check,pnpm check:packagesTIMESTAMP, andTIMESTAMP_NS. Every batch was Gorilla-flagged,mainand this PR decoded identical values throughquery()andqueryViews(), and a sample matched HTTP/exec.