Skip to content

Optimize stable full-text pagination and update Tantivy - #46

Merged
kylebernhardy merged 16 commits into
mainfrom
codex/optimize-tied-score-search
Sep 30, 2026
Merged

kylebernhardy merged 16 commits into
mainfrom
codex/optimize-tied-score-search

Conversation

@kylebernhardy

@kylebernhardy kylebernhardy commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

This improves native full-text search and indexing in five areas:

  • stable BM25 pagination for equal-score and multi-clause any searches
  • exact-total queries that reuse counts already collected by the search pass
  • smaller surface-term indexes and faster cleanup of replaced documents
  • Tantivy 0.26.2, including its buffered-union scorer correction
  • opt-in Claude and Gemini review using HarperFast's shared native-addon review scope

The public API and wire protocol are unchanged.

Reviewer focus

  1. Candidate-free, multi-clause top-level any searches use the same score-and-ID collector path for every page. Clause count is analyzed terms multiplied by selected fields, so a single-term search over two fields takes this path. This avoids cross-page score drift from mixing scorer topologies, but it enumerates every match and reads the ID fast field for each match. It remains in the cheap admission class and checks its deadline after collection. The observed one-million-record benchmark stayed below the 50 ms p99 target, but it did not isolate the bounded, non-exact shape that lost pruning. Qualification at catalog scale remains open in Prove native checkpoint publication and the Harper derived-index path.
  2. Surface companion fields now store term frequencies without positions. Existing native indexes created with both surfaceTerms: true and positions: true fail to open with E_SCHEMA_MISMATCH and must rebuild from source before serving searches. Rebuild time at catalog scale has not been measured.
  3. Writers retain Tantivy's default merge policy except for a 50% deleted-document threshold. The observed replacement workload completed faster without constraining merge-thread parallelism.
  4. A missing ID ordinal remains a fail-closed index-invariant error. Returning a partial result would make totals and pagination incorrect.
  5. The stable collector is wrapper-owned code, but it still delegates scoring, postings, top-N selection, and fast-field access to Tantivy. A direct A/B against Tantivy's native score-plus-string collector kept the wrapper collector because batched ID resolution was consistently faster. The design now calls out this ownership and requires the comparison to be repeated when either path changes.

Changes

  • Flatten any term queries into one Should union.
  • Rank stable pages by score, then raw UTF-8 ID, with deterministic segment merging.
  • Resolve result IDs and versions with one ordered dictionary traversal per string column per segment.
  • Reuse the stable collector's match count for exact totals, including boundary-tie fallback.
  • Reuse an exhausted bounded result's exact count instead of running a separate Count query.
  • Store surface companion fields with IndexRecordOption::WithFreqs.
  • Set LogMergePolicy::set_del_docs_ratio_before_merge(0.5) on every writer.
  • Update Tantivy from 0.26.1 to 0.26.2.
  • Document the Tantivy ownership boundary, schema rebuild behavior, and benchmark comparison policy.
  • Add commit-pinned Claude and Gemini PR reviewers with structural caller validation.

File guide:

Observed performance

These are local internal-harness results, not catalog-scale guarantees.

Workload Before After Result
500k varied-score exact first page, concurrency 1 ~239 QPS ~342 QPS ~43% higher throughput
500k varied-score exact first page, concurrency 8 ~929 QPS ~1,339 QPS ~44% higher throughput
1M surface-term index 52.5 MiB 43.9 MiB ~16.4% smaller
Full replacement workload 12.8 s 4.7 s ~63% lower elapsed time
10k result-metadata lookups, 100 iterations 3.31 s 51.2 ms batched dictionary traversal
100k multi-term any, concurrency 1, two alternating runs 3,015–3,065 QPS native composite 3,475–3,587 QPS wrapper collector ~13–19% higher throughput

The one-million-product query run produced:

Query Concurrency QPS p99
Filtered 1 27.09 113.0 ms
Filtered 8 59.10 352.2 ms
BM25 any 1 184.77 9.4 ms
BM25 any 8 437.36 30.9 ms

Filtered p99 remains above the proposed 50 ms target. The stable bounded any path remains the primary catalog-scale risk.

Rejected experiments

Observed locally:

  • A local score floor reduced any throughput from 3,713 to 3,605 QPS.
  • Passing that floor into Tantivy's pruning hook reduced throughput to 2,059 QPS.
  • Limiting merge threads to one reduced median ingest throughput by about 25%; limiting them to two reduced 16-index ingest throughput by about 12%.
  • A 30% deleted-document merge threshold reduced disk use further but added close-time merge drain. The 50% threshold is the conservative default for this change.

Verification

  • npm run check: formatting, TypeScript, clippy, 64 Rust tests, 142 Node tests, 5 worker tests, examples, and package checks
  • npm run test:optimized-native: optimized native concurrency
  • cargo deny check: advisories, bans, licenses, and sources
  • Pagination coverage across segments, successive pages, exact and bounded totals, past-end offsets, reopen, and tie fallback
  • Schema mismatch coverage for the former positioned surface-field layout
  • Delete-aware merge coverage for live documents and checkpoint preservation across reopen
  • Tantivy compatibility check: create with 0.26.1, commit with 0.26.2, then reopen and query with 0.26.1
  • Two alternating 100,000-document A/B runs comparing the wrapper collector with Tantivy's native composite collector
  • Documentation format and diff checks
  • Claude checked the benchmark documentation against the CI and release workflows and returned LGTM
  • Parsed all three workflow files, verified identical provider checklists and immutable pins, and confirmed every pinned reusable path exists

Claude and Cursor Grok reviewed the runtime code. Google Gemini CLI reviewed the preceding runtime state; both findings from that pass are fixed. A final local ownership audit compared the implementation with Tantivy 0.26.2 and prompted the native-collector A/B above. Claude Opus and the Gemini CLI both cleared the final AI-review workflow commit after their findings were addressed.

Related: Prove native checkpoint publication and the Harper derived-index path

Comment generated by kAIle (gpt-6-astra)

Complexity: complicated

Review-Coverage: authored=codex; ran=claude,cursor-composer,cursor-grok; adjudicated=domain; blocked=gemini(auth); rounds=11; full=5 @ 2854d90

Human-Review-Need: 4 @ 2854d90

@kylebernhardy kylebernhardy changed the title Optimize stable full-text search pagination Optimize stable full-text pagination and update Tantivy Sep 30, 2026
@kylebernhardy kylebernhardy added the claude-review Trusted-member gesture: opt this PR into Claude review. label Sep 30, 2026
@kylebernhardy
kylebernhardy merged commit 6856c8f into main Sep 30, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

claude-review Trusted-member gesture: opt this PR into Claude review.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant