Native Tantivy full-text indexing for Node.js.
This repository is under active development. The native entry point provides a standalone Tantivy
index backed by MmapDirectory. Applications own their source records and may reuse compatible
local index files or rebuild them from that authoritative data.
- Node.js 22.18 or newer, or Node.js 24 or newer
- Rust 1.90 when building from source
The package never compiles or downloads native code during installation. It installs a matching exact-version native package through npm optional dependencies.
npm install @harperfast/fulltextDo not omit optional dependencies: the matching native package is selected from them at runtime.
Applications import the single public entry point, @harperfast/fulltext/native.
The CI-qualified targets are Linux x64 and Linux arm64 with glibc 2.35 or newer, macOS arm64, and
Windows x64. Additional targets are added only after their packed artifacts are loaded and tested
on the target runtime. Unsupported targets fail with E_NATIVE_ADDON_NOT_FOUND when the native
entry point is first used; importing the JavaScript facade does not eagerly load an addon.
import { openNativeFullTextIndex } from '@harperfast/fulltext/native';
const index = await openNativeFullTextIndex({
path: './search/products',
indexId: 'products',
generation: 'v1',
fields: [{ name: 'title', weight: 3 }, { name: 'description' }],
analyzer: 'english@2',
limits: {
indexingThreads: 2,
searchThreads: 4,
writerMemoryBytes: 60_000_000,
maxQueuedCommands: 128,
maxQueuedBytes: 16 * 1024 * 1024,
maxBatchBytes: 8 * 1024 * 1024,
},
});
try {
await index.applyMutationBatch({
upserts: [{ id: 'shoe-1', version: '42', fields: { title: 'Trail running shoe', description: 'Waterproof' } }],
});
await index.commit();
await index.reload();
console.log(await index.search({ text: 'waterproof running shoes', limit: 10 }));
} finally {
await index.close({ mode: 'rollback' });
}The quick start uses standalone commit and reload. Applications that track an external checkpoint
should use publish(payload) instead, as shown below, so mutations and that checkpoint commit
together.
Every example runs against the local build and the packed npm package in CI:
| Example | Demonstrates |
|---|---|
examples/basic.mjs |
Open, mutate, commit, search, and close |
examples/query-modes.mjs |
BM25, phrase, prefix, fuzzy, and candidate filtering |
examples/checkpoint-recovery.mjs |
Atomic checkpoint publication, inspection, and reopen |
examples/highlighting.mjs |
Match tracing, UTF-16 spans, and opt-in snippets |
examples/multi-index.mjs |
Optional process limits and concurrent independent indexes |
examples/reset-rebuild.mjs |
Safe retirement, cleanup, and rebuilding a generation |
From a repository checkout, build the native addon and run any example through the local-build loader:
npm run build
node scripts/run-local.mjs examples/basic.mjsMutation batches are versioned packed values, so indexing crosses Node-API once per batch rather
than once per document. One dedicated actor owns Tantivy's single writer for each index. A bounded
search pool shares immutable searchers and can execute reads while indexing or commit work is in
progress. Queue limits reject overload with E_QUEUE_FULL rather than blocking the JavaScript
thread. applyMutationBatch() is the normal logical mutation API. It partitions one logical batch
into bounded native frames, applies them sequentially, validates native counts, and returns
{ processed, rejected, encodedBytes, frames }. IDs must be distinct across the logical batch.
By default, any rejected record fails the operation. Pass { rejectedUpsert: 'delete' } when an
unindexable replacement must delete previously searchable content for the same ID. A failure after
native application is attempted leaves the handle incomplete: writer operations reject
E_BATCH_INCOMPLETE until the handle is closed with { mode: 'rollback' }. A writer operation that
only races an in-flight logical batch rejects with E_BATCH_ACTIVE; wait for that batch to settle
and retry instead of rolling it back. This prevents a later publish from exposing part of a logical
batch. A closing or closed handle rejects every logical batch with E_CLOSED, including an empty
batch, without taking the latch.
configureNativeFullTextRuntime() is optional and idempotent for an identical configuration. It
must be called before the first successful open when one process may host many indexes; configuring
while an open is pending fails with retryable E_LOCK_BUSY; configuration after an unbudgeted open
fails with E_RESOURCE_LIMIT. A failed first open that proves its native resources were released
does not prevent later configuration. An
unproven pre-publication teardown blocks configuration until restart. The governor bounds aggregate
resident indexes, indexing and search threads, writer memory, configured queue bytes, and concurrent
expensive searches. Each writer reserves twice its maxQueuedBytes for the writer and shared search
queues; each read-only handle reserves it once for search. Every writer or reader also consumes one
resident-index slot and its configured search-thread count. A conflicting second configuration or an
open that would exceed an admission cap fails with E_RESOURCE_LIMIT; the library never evicts a
live generation. Expensive searches wait off the JavaScript thread for process capacity, bounded
by the request deadline and interrupted by close, allowing the search queues to apply backpressure.
A flexible worker serves ordinary searches while the expensive-search budget is saturated. With a
one-thread index, ordinary and expensive work share one FIFO, so a process-wide permit held by a
different index can delay ordinary work queued behind an expensive request. Closing an index
interrupts expensive requests that are waiting for a permit with E_CLOSED.
A completed close has released its reservation.
If teardown cannot prove native resources were released, the path is quarantined and its aggregate
capacity remains reserved until process restart.
Callers that omit this function retain the per-index limits.
Schema mismatches fail the whole call even with { rejectedUpsert: 'delete' }; treating schema drift
as record-local rejection could remove many documents under the wrong schema. The option applies
only to record-local E_INVALID_ARGUMENT and E_BATCH_TOO_LARGE rejections with usable IDs.
Delete mode always preflights every ID as a bounded delete frame before native admission;
assumeDistinctIds skips duplicate detection, not that feasibility scan. An unusable or oversized
ID therefore fails with nothing staged. Schema drift is detected during framing and can leave the
handle incomplete when earlier frames were already admitted.
Trusted callers that already enforce distinct IDs may pass assumeDistinctIds: true to skip the
whole-batch duplicate prepass. Supplying duplicates with that option violates the API contract.
The library snapshots the two mutation arrays, but callers must not mutate record objects, field
maps, or nested field-value arrays until the returned promise settles.
encodeMutationBatch(batch, maxBytes) rejects output beyond its encoding bound with
E_BATCH_TOO_LARGE; maxBytes defaults to 8 MiB and is intended for low-level callers producing a
single native frame. Low-level callers may use index.encodeMutationBatches(). It uses the limit from
the index's open configuration, validates IDs and fields against that handle, and greedily emits
admissible frames. IDs must be distinct across the logical batch. Aggregate size creates more
frames; a single invalid or unencodable mutation is returned in rejected with its operation and
zero-based index in the corresponding input array. It is never dropped automatically.
Encoding is synchronous and runs on the JavaScript thread. The total encoded output of one logical
call defaults to 64 MiB and can be lowered with encodeMutationBatches(batch, { maxTotalBytes }).
The default behavior throws E_BATCH_TOO_LARGE when the complete logical batch exceeds that
ceiling. Callers that pass allowPartial: true instead receive a leading prefix;
consumedUpserts and consumedDeletes identify the mutations represented by the returned frames
and rejections so they can continue with each array's remaining suffix. The low-level API validates
distinct IDs across the full input before applying the output ceiling, so repeated suffix calls
repeat that validation scan. Prefer applyMutationBatch() for large logical batches.
Producers using the low-level API should keep logical batches comfortably below that limit.
Applying multiple frames stages them in one Tantivy writer. Commit or publish only after every frame
succeeds; on a later failure, close with rollback rather than publishing the partial logical batch.
Concurrent low-level callers should leave queue-byte headroom for commit or publish.
A successful apply() resolves to the number of accepted mutation commands, including deletes for
IDs that are not currently indexed.
One caller must own a handle's complete apply-and-publish sequence at a time. The high-level method blocks low-level writer interleaving while its logical batch is active. The low-level frame API does not infer logical batch boundaries, so callers using it remain responsible for excluding another publisher between frames. Searches may still run concurrently with either writer sequence.
Search uses weighted BM25 and a typed request; raw Tantivy query syntax is not exposed. mode is
any, all, phrase, prefix, fuzzy, or fuzzy-prefix. The older operator: 'any' | 'all'
spelling remains a compatibility alias and cannot be combined with mode. Candidate IDs compile to
a score-neutral required filter. Prefix modes are autocomplete-oriented, require offset zero, and
return at most 100 hits. Fuzzy and prefix work has fixed term, clause, expansion, request, result,
and execution ceilings. Prefix expansion fails with E_PREFIX_TOO_BROAD instead of silently
truncating the term set and returning incomplete rankings. fuzzy-prefix is a preview capability
until catalog-scale benchmark qualification is complete.
Record IDs are limited to 4,096 UTF-8 bytes, keeping deterministic tie-page sorting memory bounded.
An upsert may include an opaque version string of up to 4,096 UTF-8 bytes. The version is returned
with its search hit, allowing a derived-index consumer to discard a hit when the authoritative
record has moved past the indexed version.
Use query for Boolean expressions within one index. Expressions may nest to eight levels and are
bounded by the same request, term, and clause limits as simple searches:
const result = await index.search({
query: {
operator: 'and',
clauses: [
{ text: 'waterproof trail', mode: 'all', fields: ['title'] },
{ operator: 'not', clause: { text: 'used' } },
],
},
});Negation filters do not add to BM25 scores. A query made only of negation matches assigns every surviving hit a score of zero and orders ties by UTF-8 ID.
total is bounded by default so Tantivy can retain block-max WAND pruning where the query shape
supports it. Set exactTotal: true only when an exact match count is required. A separate count
pass runs only when the result was not exhausted and the selected collector did not already visit
and count every match. Ranking is score descending, then UTF-8 ID ascending, including ties that
cross segment or page boundaries.
Candidate-free, top-level any searches with more than one scoring clause use the same stable
score-and-ID collector path for every page. Clause count is analyzed terms multiplied by selected
fields, so one term searched across two fields takes this path. This prevents segment merges and
different page boundaries from changing tie order, but it visits every match and checks the request
deadline after collection. Other bounded query shapes start with Tantivy's score-pruned TopDocs
path; a score tie at the requested page boundary falls back to the stable all-match collector.
positions defaults on and is required for phrase search. surfaceTerms defaults off and creates
an internal unstemmed companion term field used by prefix, fuzzy-prefix, and match tracing. It does
not store source values or positions; it retains term frequencies for scoring. Both settings are
persisted and must match when the index is reopened. Indexes created before this change with both
surfaceTerms: true and positions: true fail to open with E_SCHEMA_MISMATCH and must be rebuilt
from their authoritative source. Enable surfaceTerms only on indexes that need those operations.
Field weights are query-time boosts and may change on reopen without rebuilding the index.
english@2 applies Unicode NFKC normalization, lowercase and Latin-to-ASCII folding, English
possessive removal, optional English stop words, and English stemming. Token offsets continue to
refer to the original source value. Normalization is streamed with source-span tracking, and token
text is bounded before Tantivy's long-token filter, avoiding memory growth proportional to Unicode
compatibility expansion. Inputs that exceed Unicode's 30-non-starter Stream-Safe limit receive a
standard combining-grapheme-joiner boundary before normalization. Index-time synonyms are optional
and bounded. Each source and replacement must produce exactly one normalized and analyzed term.
Rules are canonicalized,
persisted in the index identity, and expanded once at the source token's position; query text is not
synonym-expanded. The expansion is also written to the surface field, so prefix autocomplete can
match replacement terms. Tantivy counts the stacked alternatives in BM25 field length, which can
lower unrelated-term scores for documents containing a synonym source. Changing analyzer settings
or synonyms requires a new generation/rebuild. Match tracing applies the same index-time expansion
so synonym-derived hits map back to the original token span.
The public analyzer name and identity-sidecar version jointly identify persisted analyzer semantics. Any future filter or ordering change must use a new analyzer name or sidecar version so existing indexes fail closed and rebuild instead of silently changing recall.
Highlighting is opt-in and operates on caller-supplied current source values, so the library never returns stale stored text. It returns UTF-16 half-open offsets and no HTML. Snippets are also off by default:
const traced = await index.traceMatches(
{ text: 'trail running', mode: 'phrase' },
[{ id: 'shoe-1', fields: { title: 'Waterproof trail running shoe' } }],
{ snippets: true, fragmentLength: 160, maxFragmentsPerValue: 3 },
);Tracing evaluates only the supplied current values; it does not expand the live index dictionary.
Analysis is limited to 262,144 emitted tokens per value after synonym expansion. When that token
ceiling, a span ceiling, or the response ceiling is reached, complete is false and later matches
may be omitted.
Search and tracing share a maximum 30-second queue-plus-execution budget. Applications can pass a
shorter remaining request budget through the second method argument; larger values are clamped to
30 seconds. The budget determines whether queued or completed work is accepted; it is not a
callback timer. If every eligible worker is already executing non-interruptible work, an expired
queued request is rejected when a worker next examines it. Tantivy search itself is not
interruptible, so a search that expires in flight is discarded after Tantivy returns. Match tracing
checks its deadline while tokenizing and matching. With at least two search threads, one worker is
reserved for ordinary bounded any/all BM25. The remaining workers prioritize phrase, prefix,
fuzzy, unanchored-negation, exact-total, and trace work, then steal ordinary work when that queue is
idle. A negation intersected with a positive ordinary clause stays in the ordinary lane. A
one-thread configuration remains valid but cannot isolate query classes.
close() rejects uncommitted data by default. Use close({ mode: 'rollback' }) to discard it
explicitly. commit() publishes mutations, and reload() makes the latest commit visible to this
handle's searches. A reload that cannot align Tantivy's snapshot with its checkpoint after three
bounded attempts returns E_RELOAD_FAILED; callers can retry because the reader remains open.
The failed attempt leaves the reader on its previous aligned snapshot and checkpoint.
One process may open one writer and multiple readers for the same physical index. Readers share the same physical files, but each has an independently bounded search runtime and never reserves Tantivy's writer lock:
import { openNativeFullTextReader } from '@harperfast/fulltext/native';
const reader = await openNativeFullTextReader(options);
await reader.reload(); // call after the writer publishes a newer checkpoint
const result = await reader.search({ text: 'trail shoe' });
await reader.close();Readers use Tantivy's manual reload policy. They do not watch or poll the filesystem; the caller
coordinates publication and calls reload() only when a newer revision is available. Reset rejects
while any writer or reader is live in the current process; separate processes must coordinate reset
with their own reader lifecycle. A reader never creates missing storage and fails with
E_INDEX_NOT_READY until a writer has created a complete index. A successful reload also refreshes
the reader's committedPayload.
Use publish(payload) when a consumer needs to resume from a durable checkpoint:
console.log(index.committedPayload); // undefined on a new index; recovered from files on reopen
await index.applyMutationBatch({
upserts: [{ id: 'shoe-1', fields: { title: 'Trail shoes' } }],
});
await index.publish('source-checkpoint-42');
// Both the mutations and the checkpoint are committed; searches now see that commit.The payload is an opaque, well-formed Unicode string, limited to 64 KiB in UTF-8. Fulltext does not
interpret it or verify that it describes the mutations supplied by the caller. Empty strings are
valid checkpoints; undefined means the index has never committed one. Publication uses the same
bounded writer queue as mutations and returns a Tantivy opstamp, not an application cursor.
Once a generation has a checkpoint, plain commit() rejects with E_CHECKPOINT_REQUIRED, including
after reopen. Pending mutations remain available to a subsequent publish(), or can be discarded
with rollback close. A publication without mutations can advance the checkpoint. Ordinary indexes
that never publish retain the separate commit() and reload() API.
An accepted publication that fails can have an ambiguous durable outcome: the commit may have
succeeded before reader reload failed. The handle becomes terminal and committedPayload throws
E_POISONED; close it, reopen the same files and identity, and recover the checkpoint from that new
handle. Do not infer recovery progress from status() counters. Validation and queue admission
failures do not themselves poison the handle or invalidate a known checkpoint.
inspectNativeFullTextIndex(options) synchronously validates an existing index's identity, schema,
metadata, and committed payload without creating files, reserving a handle, starting actors, or
acquiring the Tantivy writer. It returns missing, cursorless, checkpointed with the committed
payload, or incompatible with a stable identity/schema or corrupt-format code. checkpointed
means the committed metadata is compatible; callers must still open the index successfully before
serving queries. Opening maps missing segment files and unreadable segment metadata or footers to
rebuildable error codes.
Operational storage failures throw. Inspection is intended for short lifecycle checks such as
derived-index election, not request hot paths. Its options intentionally omit writer, queue, and
search limits because inspection creates none of those resources.
validateNativeFullTextIndexOptions(options) runs the same native configuration decoder as open
without creating files or starting actors. This is intended for activation-time validation.
close() is also the native quiescence barrier. A resolved result means the writer, search actors,
readers, merge threads, and memory mappings no longer use the index path. After closing an
incompatible or corrupt local index, retire it atomically before rebuilding:
const result = await resetNativeFullTextIndex({ path: './search/products', indexId: 'products' });
if (result.state === 'reset') {
console.log(`retired native files at ${result.retiredPath}`);
}The resolved close result is {} normally. If native resources were released but shutdown also
reported an operational cleanup error, it is { cleanupError } and that error has code
E_CLOSE_FAILED; the path is still safe to reset. E_QUIESCENCE_FAILED rejects because the library
could not prove all native work stopped. Do not reset, remove, or rename that path until the process
restarts.
Reset returns missing without creating the path. It rejects a live owner with E_LOCK_BUSY, a
different persisted logical index with E_IDENTITY_MISMATCH, and unrelated nonempty directories
with E_INVALID_ARGUMENT. A malformed identity sidecar fails closed with E_INDEX_CORRUPT. On
success, reset renames the live directory into a unique path below the parent's .fulltext-retired
directory. The library does not delete it automatically. Call
reclaimRetiredNativeFullTextIndexes({ path, retiredPath: result.retiredPath }) at a lifecycle point
chosen by the application. Passing the opaque reset result verifies that the hint belongs to the
requested index. Use the same path value for reset and reclaim so generated names match. The
reclaimer removes only retired trees generated for that index path, ignores unrelated entries and
other indices, and returns { removed, failed }. Applications choose when retired files are safe
to reclaim.
An established duplicate open returns E_DUPLICATE_OPEN. An open racing another open or reset can
return E_LOCK_BUSY while the shared lifecycle lock is held; callers may retry that acquisition.
An open that cannot capture one checkpoint-aligned Tantivy snapshot returns E_RELOAD_FAILED; retry
the open.
Lifecycle lock files are stored in the index parent's .fulltext-locks directory so reset can keep
the handoff lock while renaming the native directory on Windows. The library does not remove this
lock directory. The parent must permit creating this directory, and .fulltext-locks must remain
writable while indices are opened or reset; failures name the lock-directory path.
This API uses native ABI 8 and packed protocol 4. The loader rejects older addon binaries. Native
identity sidecar v4 fingerprints the internal source-version field in addition to canonical
synonyms and the completed english@2 semantics. Older v2 and v3 indexes remain identifiable for
safe reset but cannot be reopened under the new schema; rebuild them from the authoritative source.
All public failures are FulltextError instances with a stable code. Branch on the code instead
of parsing the message:
import { FulltextError, openNativeFullTextIndex, runtimeInfo } from '@harperfast/fulltext/native';
console.log(await runtimeInfo());
try {
await openNativeFullTextIndex(options);
} catch (error) {
if (error instanceof FulltextError && error.code === 'E_IDENTITY_MISMATCH') {
// Retire and rebuild this derived index generation.
} else {
throw error;
}
}runtimeInfo() reports the package, Tantivy, ABI, query API, lifecycle API, mutation API, supported
storage backend, and hard protocol limits. It does not open an index or create files.
The package has one public entry point: @harperfast/fulltext/native. It stores application-owned
derived indexes in Tantivy's native filesystem directory. The addon does not link RocksDB, depend
on rocksdb-js, or fall back to another storage backend.
npm ci --ignore-scripts
npm run build:debug
npm test
npm run lint
npm run format:check
npm run benchmark:native -- --documents 100000 --concurrency 4 --commit-every 25000
npm run benchmark:multi-index -- --indexes 8 --documents 100000 --concurrency 16
npm run benchmark:native -- --documents 100000 --revision v0.1.0 --output benchmark-native.json
npm run benchmark:native -- --documents 100000 --concurrency 4 --commit-every 25000 --mutation-driver low-level
npm run benchmark:inspect -- --indexes 1,10,100,1000 --commits 64 --warm-rounds 10CI reports JavaScript/TypeScript and Rust coverage separately. The Node report uses the built-in test runner coverage and stores raw V8 coverage; the Rust report stores LCOV output. Coverage is reported for visibility and release-to-release comparison, not enforced as a single blended gate.
The benchmark generates a deterministic, high-cardinality product catalog and emits one versioned
JSON record. It reports packing, apply, durable end-to-end ingestion, actor queue and execution
time, logical-batch p50/p95/p99, frame count, commit distributions, reload cost, warm and cold query
p50/p95/p99, per-mode term/phrase/prefix/fuzzy/candidate performance, exact-total overhead, index
bytes, and periodically sampled process RSS. The default
--mutation-driver logical exercises applyMutationBatch(); rerun the identical command with
--mutation-driver low-level for the prior encode-plus-apply path. Compare
mutationDriverMilliseconds, durable end-to-end throughput, logical-batch percentiles, and peak
RSS. The component packingMilliseconds and applyMilliseconds fields are not comparable because
the logical API performs both inside one call. --commit-every sets the target number of mutations
between durability points; it materially affects throughput and peak memory because
replacement-safe upserts include delete terms. CI runs only the correctness smoke profile; timing
comparisons require controlled hardware.
benchmark:multi-index exercises independent writers in parallel under the process budget. It
includes heavy-tail field sizes, bounded synonyms, update/delete churn, resource-cap rejection,
warm concurrent queries, close/reopen, and cold queries. CI runs the two-index smoke profile; the
release workflow records the larger eight-index profile per architecture. These results measure
the library's native path, not source projection, authorization, or record retrieval performed by
an application.
--revision labels a result and --output writes the same JSON record printed to stdout. Pull
requests keep smoke records as GitHub Actions artifacts for 30 days. After a GitHub release is
published, the release benchmark keeps 90-day Actions artifacts and attaches native and multi-index
JSON results for Linux x64 and Linux arm64 to that release. These assets provide a permanent
per-architecture release-over-release history, but the post-publish workflow does not gate npm
publication, calculate a baseline delta, or fail a build on timing. Shared-runner numbers are
evidence that the workload still runs, not a latency gate. Compare performance only with the same
benchmark format, workload, architecture, and controlled hardware. Record observed numbers in the
pull request or release notes; keep this README focused on the reproducible method rather than
environment-specific targets.
The inspection benchmark compares synchronous read-only inspection with full writer reopen across
multiple index counts. It reports equivalent first-pass and warm p50/p95/p99/max latency,
synchronous wall time per inspection sweep, metadata size, and the actual segment count produced by
the seed workload. It does not claim to model OS-cold storage. --indexes accepts configurations
that create at most 10,000 temporary index directories across both benchmark paths.
Generated Node-API declarations in ts/addon.d.ts are private implementation types. Consumers use
only the types exported from a package entry point.
The root @harperfast/fulltext package contains only the JavaScript and TypeScript facade. Native
artifacts are published as exact-version platform packages and selected at runtime. Release tags
must match both npm and Cargo versions. The release workflow builds and tests each supported native
artifact, installs the packed root and platform tarballs in a clean consumer with lifecycle scripts
disabled, and publishes with npm provenance. Platform packages are published before the root, so a
root version is never available until every supported artifact has passed its consumer test.
See CONTRIBUTING.md for the development workflow and the native backend design for implementation details.
Apache-2.0