Skip to content

perf(qwp): improved Gorilla timestamp decoding - #69

Open
glasstiger wants to merge 4 commits into
mainfrom
ia_gorilla_decoding
Open

glasstiger wants to merge 4 commits into
mainfrom
ia_gorilla_decoding

Conversation

@glasstiger

@glasstiger glasstiger commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Gorilla-encoded TIMESTAMP, TIMESTAMP_NANOS, and DATE columns in query results now decode without BigInt arithmetic, in packages/client-core/src/_qwp/_core/result-batch.ts. In testing, a QuestDB nightly server flagged every result batch as Gorilla-encoded, so this path decodes the timestamp columns of typical query results in both @questdb/nodejs-client and @questdb/browser-client. No public API change, and decoded values and errors are unchanged.

  • One decoder for both paths. decode() and decodeView() each held a copy of a loop that read the bitstream one bit at a time and built every delta-of-delta from per-bit BigInt operations (value |= 1n << BigInt(bit)). Both now call decodeGorillaTimestamps(), which reads through a 32-bit bit buffer, carries the int64 delta and timestamp as int32 halves with explicit carries (wrapping exactly as BigInt.asIntN(64, ...) does), and emits runs of repeated deltas in one step.
  • Little-endian output on every host. Values are stored as int32 word pairs through an Int32Array, which ran the decode loop 1.8-4.7x faster than DataView stores in V8. A byte-swap pass runs only on big-endian hosts.
  • Earlier truncation check. A stream too short to hold one bit per row is rejected before any storage is sized for it; the view path used to allocate first. The error is unchanged (truncated QWP Gorilla bitstream).
  • Retained decode buffer for decode(). One buffer per decoder is reused across batches, as column views already did. Its BigInt read-out loop is deliberately separate from readInt64Values(): with one loop serving both, V8 ran this path up to 2x slower, depending on which caller it optimized first.
  • Tests: every delta-of-delta width at both bucket edges, NULLs, int64 edges and low-word carries, and decoder reuse, through decode() and decodeView(); 600 random streams (including truncated ones and int64 wraparound) against a bit-at-a-time BigInt reference decoder, comparing values, consumed bytes, and errors; every malformed-column error through both paths.
  • Benchmarks: timestamp-only Gorilla arms in benchmarks/egress.bench.ts (constant, jittered, and irregular intervals). benchmarks/workloads.ts exports its existing xorshift generator for them.

Performance

Diagnostic, not CI gates. Before is main; two interleaved runs each, averaged.

Node 24, benchmarks/egress.bench.ts, one Gorilla TIMESTAMP column of 10k rows, batches/s:

Timestamps column view: main this PR materialized: main this PR
constant 1 ms 19,079 61,893 (3.2x) 13,142 23,208 (1.8x)
1 ms, ±50 µs jitter 915 14,340 (15.7x) 939 10,175 (10.8x)
irregular, mean 1 s 272 8,819 (32.5x) 261 7,144 (27.4x)

The existing mixed batch (INT, DOUBLE, VARCHAR, constant-interval Gorilla TIMESTAMP): reusable column views ~13.5k → ~35.7k batches/s (2.6x), decode() materialized 861 → 897 (1.04x); the traversal and Zstd arms do not exercise this code and are unchanged within noise.

JavaScriptCore (Bun 1.0.14, Safari's engine), the built browser bundle, the same three interval shapes, 10k rows, median µs per batch:

Timestamps column view: main this PR materialized: main this PR
constant 1 ms 896 12.4 (72x) 520 186.5 (2.8x)
1 ms, ±50 µs jitter 4,049 61.5 (66x) 6,491 253.7 (26x)
irregular, mean 1 s 11,888 110.8 (107x) 17,782 333.8 (53x)

The gain is larger in JavaScriptCore because it allocates on every BigInt operation, while V8 lowers BigInt.asIntN(64, ...) arithmetic to native int64 math. Materializing a bigint[] still costs one BigInt allocation per value, which bounds the materialized speedup.

Validation

  • pnpm exec vitest run (1,126 tests, including container integration), pnpm test:dist, pnpm test:qwp-browser
  • pnpm typecheck, pnpm typecheck:test, pnpm typecheck:qwp-browser, pnpm typecheck:bench, pnpm typecheck:dist
  • pnpm eslint, pnpm lint:bench, pnpm format:check, pnpm check:packages
  • Mutation check: 18 decoder faults injected one at a time (dropped carries, lost sign extension, an unclamped zero run, a short refill, removed or off-by-two truncation checks, rounded-down consumed bytes, wrong output lengths, a byte swap on a little-endian host). The new tests fail for each.
  • Against a QuestDB 10.0.2-SNAPSHOT server (nightly image): 300k-row queries over constant, jittered, and irregular designated timestamps, an unordered non-designated TIMESTAMP, and TIMESTAMP_NS. Every batch was Gorilla-flagged, main and this PR decoded identical values through query() and queryViews(), and a sample matched HTTP /exec.
  • Not exercised: the big-endian byte-swap path (no big-endian host available) and Firefox.

decode() and decodeView() each held a copy of a Gorilla loop that read
the bitstream one bit at a time and built every delta-of-delta from
per-bit BigInt operations. Both now call one decoder that reads through a
32-bit bit buffer, carries the int64 delta and timestamp as int32 halves
with explicit carries, wrapping exactly as BigInt.asIntN(64, ...) does,
and emits runs of repeated deltas in one step. Values are stored as int32
word pairs through an Int32Array, with a byte swap on big-endian hosts so
the bytes stay little-endian. Decoded values and errors are unchanged.

A stream too short to hold one bit per row is now rejected before any
storage is sized for it, and decode() reuses one decode buffer per
decoder across batches, as column views already did.

On 10,000-row batches in Node 24, column views decode 3.2-32x faster and
materialized batches 1.8-27x faster, least for a constant interval and
most for irregular timestamps. JavaScriptCore views decode 66-107x faster.

Adds code-width, NULL, int64-edge, and decoder-reuse tests for both
paths, a 600-stream comparison against a bit-at-a-time BigInt reference
decoder, every malformed-column error, and timestamp-only Gorilla arms in
benchmarks/egress.bench.ts.
The code-width test reused one decoder only with shorter batches, so
nothing exercised growing the retained Gorilla storage: a decoder that
stopped growing it passed every test while a larger later batch threw a
RangeError. A three-value batch now precedes the main frame, which has
to grow that storage in both decode() and decodeView().

The test's `max` column claimed to cover the top of the int64 range but
peaked about 7.1e11 below INT64_MAX. It now ends exactly at INT64_MAX.
The Gorilla decoder writes int32 words through an Int32Array and
byte-swaps them on big-endian hosts, so column views still expose
little-endian int64 bytes. No test reached that branch:
NATIVE_LITTLE_ENDIAN is fixed when result-batch.ts loads and every CI
host is little-endian, so deleting the swap call left all 897 QWP tests
passing.

The new test makes the module's Uint16Array.of(1) probe see big-endian
storage, loads a fresh copy, and decodes through decode() and
decodeView(). On a little-endian host the swap then shows as
byte-reversed int32 halves in every value. It fails if the swap call is
deleted or inverted, or if the swap drops a byte, shifts one wrong, or
covers only half the words.
The decoder's comment said the speedup was least for a constant interval
and most for irregular timestamps, in V8 and JavaScriptCore alike. In
JavaScriptCore the jittered series gains least: two Bun 1.0.14 runs
measured 66x and 61.7x for it, against 72x and 72.5x for a constant
interval. The ordering now applies to V8 only, and the JavaScriptCore
figure reads "over 60x" rather than a 65-110x range that the second run
fell below.
@glasstiger

Copy link
Copy Markdown
Collaborator Author

Level 3 review — approve (head cc1dc99766c1d69128b7a63f60919d1a896ad678).

No admitted findings: Critical 0, Moderate 0, Minor 0; in-diff 0, out-of-diff breakage 0. Test gate passes with 0 admitted coverage gaps. The full suite passed (1,127 tests), along with package-boundary and browser suites, typechecks, lint, and formatting checks. Base/head benchmarks reproduced the claimed Gorilla decoding speedup. The decoder retains row-cap-bounded scratch storage; no regression was observed.

Submodules: none.

@glasstiger

Copy link
Copy Markdown
Collaborator Author

Level 3 review — approve (head cc1dc99, base a658130)

No findings in this PR: Critical 0, Moderate 0, Minor 0. The test gate passes with no coverage gaps worth reporting.

Evidence

  • Values and errors unchanged. About 160k random Gorilla-flagged RESULT_BATCH frames were decoded by both base and head builds, on Node 24 and on Bun 1.0.14 (JavaScriptCore). The frames covered NULL bitmaps, raw fallback, unknown encodings, truncated and garbage streams, int64 edges, and decoder reuse. decode() values, decodeView() values and valuesBytes(), materialize(), and error class and message were identical in every case: about 57.5k frames decoded, and the rest failed with the same error on both.

  • Speedups reproduce. Node 24, 10k rows, base → head:

    Path constant jittered irregular
    column views 3.1x 15.2x 32.1x
    decode() 1.8x 10.9x 26.3x

    On Bun (JavaScriptCore), column views were 79–173x faster.

  • Tests. Every injected decoder fault that changed decoded values or the bytes consumed was caught by the new tests.

  • Checks at head, all passing: every tsc typecheck including the built .d.ts files (also under TypeScript 4.9), eslint, lint:bench, prettier, vitest test/qwp (898 tests), test:dist, check:packages, and test:qwp-browser (Chromium).

  • Not run: the container integration suite (it covers ILP ingress, not this decoder), Firefox, and a real big-endian host.

Tradeoff: decode() now keeps one scratch buffer per decoder: 128 KiB after a 10k-row column, at most 8 MiB at the 1,048,576-row cap, and lower with maxBatchRows. The PR states this, and it follows the existing view-path scratch.

Pre-existing bug (not from this PR, not blocking; worth its own issue)

QwpResultBatchDecoder.resetQuerySchema() throws a TypeError once decodeView() has used a slot above a never-used lower slot. The loop at result-batch.ts:1578 calls release() on an array hole, while releaseView() at line 1618 already guards with ?.. Egress sessions hand out slots in order from 0, so only direct users of the public decoder can hit it.

const decoder = new QwpResultBatchDecoder();
decoder.decodeView(message, 1); // slot 0 never used
decoder.resetQuerySchema(); // TypeError: Cannot read properties of undefined (reading 'release')

The schema and batch-sequence reset after the loop then never run. Reproduced on base and head (Node 24, Bun 1.0.14). Suggested fix: batch?.release(), plus a public-API test that decodes into slot 1 and then resets.

Submodules: none.

@glasstiger glasstiger changed the title perf(qwp): decode Gorilla timestamps without BigInt perf(qwp): improve efficiency of Gorilla timestamps decoding Oct 2, 2026
@glasstiger glasstiger changed the title perf(qwp): improve efficiency of Gorilla timestamps decoding perf(qwp): improved Gorilla timestamps decoding Oct 2, 2026
@glasstiger glasstiger changed the title perf(qwp): improved Gorilla timestamps decoding perf(qwp): improved Gorilla timestamp decoding Oct 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant