API Reference
Index
The primary entry point. Wraps the Laurus search engine.
class Index {
static create(
path?: string | null,
schema?: Schema,
walSyncPolicy?: WalSyncPolicy,
commitPolicy?: CommitPolicy,
): Promise<Index>;
}
Factory method
| Parameter | Type | Default | Description |
|---|---|---|---|
path | string | null | null | Directory for persistent storage. null creates an in-memory index. When given, the directory follows the <path>/schema.toml + <path>/store/ layout laurus-cli create index/--index-dir uses, so an index built here can also be opened by the CLI (and vice versa). See below. |
schema | Schema | empty | Schema definition. Only meaningful when creating a new file-backed index (or an in-memory one); must be omitted when reopening an existing file-backed index – the persisted schema is loaded instead. An empty schema is used when omitted for a new index. |
walSyncPolicy | WalSyncPolicy | per-record | WAL durability policy. Omit to keep the default per-record fsync. See WAL sync policy / durability. |
commitPolicy | CommitPolicy | manual | Auto-commit policy. Omit to keep the default manual policy (the caller drives every commit()). See Commit policy / auto-commit. |
Creating vs. reopening a file-backed index (path given): if <path>/schema.toml does not yet exist, this call creates a new index and persists schema (or an empty schema, if omitted) to it. If <path>/schema.toml already exists, this call reopens the index – schema must be omitted, or the call throws (it would be ambiguous which schema should win). It also throws if path contains an index in the layout used before this convention was introduced (segment files directly under path, no schema.toml).
Methods
| Method | Description |
|---|---|
putDocument(id, doc) | Upsert a document. Replaces all existing versions. |
addDocument(id, doc) | Append a document chunk without removing existing versions. |
putDocuments(docs) | Batched upsert. docs is an Array<[id, doc]> applied in order with one WAL fsync per batch (duplicate ids dedup, last wins). Fails fast at the first bad entry; the applied prefix is not rolled back. |
addDocuments(docs) | Batched chunk append. Like putDocuments but repeated ids accumulate as separate versions. |
getDocuments(id) | Return all stored versions for the given ID. |
deleteDocuments(id) | Delete all versions for the given ID. |
commit() | Flush writes and make pending changes searchable. |
flushWal() | Force a durable WAL barrier. See WAL sync policy / durability. |
search(query, limit?, offset?, highlight?, rescore?) | Search with a DSL string. rescore reorders the top results by late interaction; see Late-interaction rescore. |
searchTerm(field, term, limit?, offset?, highlight?) | Search with an exact term match. |
searchVector(field, vector, limit?, offset?) | Search with a pre-computed vector. |
searchVectorText(field, text, limit?, offset?) | Search with text (auto-embedded). |
searchWithRequest(request) | Search with a SearchRequest. |
searchBatch(queries, limit?, offset?, highlight?) | Execute multiple DSL string queries in parallel. results[i] corresponds to queries[i]. Returns Promise<Array<Array<JsSearchResult>>>. Empty input returns []. highlight applies identically to every query. |
stats() | Return index statistics (documentCount, vectorFields). |
All document methods and search methods are async
and return Promises. stats() is synchronous.
stats() returns an object shaped like:
interface IndexStats {
documentCount: number;
vectorFields: Record<string, { count: number; dimension: number }>;
}
WAL sync policy / durability
For a persistent index, every write is appended to a write-ahead log (WAL).
By default the WAL is fsync-ed on every record, so each write is fully
durable as soon as its Promise resolves. Index.create accepts an optional
walSyncPolicy to trade some durability for higher write throughput, and
flushWal() forces a durable barrier on demand.
class WalSyncPolicy {
static perRecord(): WalSyncPolicy;
static group(
maxRecords?: number,
maxBytes?: number,
maxIntervalMs?: number,
): WalSyncPolicy;
}
| Constructor | Description |
|---|---|
WalSyncPolicy.perRecord() | Default. fsync after every WAL record; fully durable per write. |
WalSyncPolicy.group(...) | Group commit. Batches fsync across writes. |
group(...) parameters (omit any argument to keep its default):
| Parameter | Default | Description |
|---|---|---|
maxRecords | 1024 | Flush once this many records have accumulated. |
maxBytes | 1048576 (1 MiB) | Flush once this many unsynced bytes have accumulated. |
maxIntervalMs | none | Optional periodic flush timer (milliseconds). Omit to disable the timer. |
With group commit the WAL is flushed when either maxRecords or
maxBytes is reached, and always at commit(). A crash can lose up to the
last unsynced batch — the same trade-off as SQLite’s synchronous = NORMAL.
Call flushWal() to force everything written so far to disk without a full
commit().
| Method | Description |
|---|---|
flushWal() | Force a durable WAL barrier now. Returns Promise<void>. |
import { Index, WalSyncPolicy } from "laurus-nodejs";
// Opt into group commit with a 1-second periodic flush timer.
const policy = WalSyncPolicy.group(4096, undefined, 1000);
const index = await Index.create("./myindex", schema, policy);
for (let i = 0; i < 10000; i++) {
await index.putDocument(`doc${i}`, { title: `Document ${i}` });
}
// Force a durable barrier without committing yet.
await index.flushWal();
await index.commit(); // also flushes the WAL
Omit walSyncPolicy (or pass WalSyncPolicy.perRecord()) to keep the
default, fully durable behaviour.
Commit policy / auto-commit
By default the caller drives every commit(), so pending changes become
searchable only when you call it explicitly. Index.create accepts an
optional commitPolicy to let the engine auto-commit on an ingestion-driven
cadence instead, so writes are materialized without an explicit commit().
class CommitPolicy {
static manual(): CommitPolicy;
static everyDocs(n: number): CommitPolicy;
static intervalMs(ms: number): CommitPolicy;
}
| Constructor | Description |
|---|---|
CommitPolicy.manual() | Default. No auto-commit; the caller drives every commit(). |
CommitPolicy.everyDocs(n) | Auto-commit after every n applied documents. |
CommitPolicy.intervalMs(ms) | Auto-commit at least every ms milliseconds via a background timer (default: none). Native-only; a no-op on wasm (no background threads). |
With everyDocs(n) the engine commits after every n applied documents,
counted across both singular and batch ingest — including every n
documents within a single batch. Passing everyDocs(0) is valid and
disables auto-commit, which is equivalent to CommitPolicy.manual().
intervalMs(ms) is the time-based counterpart of everyDocs(n): a background
timer auto-commits at least every ms milliseconds, so a trailing partial
batch is materialized even while ingestion is idle. It is native-only — on
wasm there are no background threads, so the engine treats intervalMs as a
no-op (the factory still constructs the value, but no timed commit happens
under WebAssembly).
commitPolicy is orthogonal to walSyncPolicy: walSyncPolicy governs WAL
fsync durability, while commitPolicy governs when the stores
materialize pending changes into a searchable commit. You can combine them
freely.
import { Index, CommitPolicy } from "laurus-nodejs";
// Auto-commit after every 1000 applied documents.
const index = await Index.create(
null,
schema,
undefined,
CommitPolicy.everyDocs(1000),
);
for (let i = 0; i < 10000; i++) {
await index.putDocument(`doc${i}`, { title: `Document ${i}` });
}
// The engine has already committed ten times; no explicit commit() needed.
Omit commitPolicy (or pass CommitPolicy.manual()) to keep the default,
caller-driven commit behaviour.
Schema
Defines the fields and index types for an Index.
class Schema {
constructor();
}
Field methods
| Method | Description |
|---|---|
addTextField(name, stored?, indexed?, termVectors?, docValues?, analyzer?, multiValued?, positionIncrementGap?) | Full-text field (inverted index, BM25). docValues controls whether the value is also copied into DocValues, the column-oriented store sort/facet/aggregation read from (Issue #1047, default true); takes effect only when stored is also true. Pass multiValued: true to accept an array of strings (Issue #1175): a term query matches if any element contains the term, a phrase query never spans two elements unless its slop reaches positionIncrementGap (default 100; 0 numbers the elements as if concatenated), and values are read back as an array of strings. analyzer is the name of a parameter-less built-in ("standard", "english", "keyword", "simple", "noop") or any custom name registered via addAnalyzer. For the parameterised Japanese preset (which requires a Lindera dictionary path), register a custom analyzer with a lindera tokenizer and reference it by name. |
addIntegerField(name, stored?, indexed?, multiValued?, docValues?) | 64-bit integer field. Pass multiValued: true to accept arrays of integers (range queries match if any value satisfies the predicate). See docValues above. |
addFloatField(name, stored?, indexed?, multiValued?, docValues?) | 64-bit float field. Pass multiValued: true to accept arrays of floats (range queries match if any value satisfies the predicate). See docValues above. |
addBooleanField(name, stored?, indexed?, multiValued?, docValues?) | Boolean field. Pass multiValued: true to accept an array of booleans (a term query such as flags:true matches if any element equals the value; values are read back as an array of booleans). See docValues above. |
addBytesField(name, stored?, multiValued?) | Raw bytes field. No docValues option: a Bytes value is never written to DocValues regardless. Pass multiValued: true to accept an array of base64 strings (Issue #1176), decoded element-wise by the schema-aware coercion the same way a single base64 string is; Bytes is never indexed, so unlike every other multiValued option this has no query-matching semantics — it only governs the stored shape and ingestion arity. Values are read back as an array of byte-integer arrays with MIME dropped, the same input/output asymmetry the scalar field already has. |
addGeoField(name, stored?, indexed?, multiValued?, docValues?) | Geographic coordinate field. Pass multiValued: true to accept an array of { lat, lon } objects (distance / bounding-box queries match if any point satisfies the predicate). See docValues above. |
addGeo3dField(name, stored?, indexed?, multiValued?, docValues?) | 3D ECEF Cartesian point field (x, y, z in metres). Pass multiValued: true to accept an array of { x, y, z } objects (distance / bounding-box / nearest queries match if any point satisfies the predicate). See Geo3d concepts and docValues above. |
addDatetimeField(name, stored?, indexed?, multiValued?, docValues?) | UTC datetime field. Pass multiValued: true to accept an array of RFC 3339 strings (range queries match if any instant satisfies the predicate; values are read back as an array of RFC 3339 strings normalized to UTC). See docValues above. |
addHnswField(name, dimension, distance?, m?, efConstruction?, defaultEfSearch?, embedder?, quantizer?, subvectorCount?, rerankStorage?, pqCodebookPath?, baseWeight?) | HNSW vector field. baseWeight sets this field’s relative scoring priority when searched alongside other vector fields (Issue #1084); see Vector Search → Weights. |
addFlatField(name, dimension, distance?, embedder?, baseWeight?) | Flat (brute-force) vector field. |
addIvfField(name, dimension, distance?, nClusters?, nProbe?, embedder?, baseWeight?) | IVF vector field. |
addMultiVectorField(name, dimension, distance?, embedder?, storage?) | Multi-vector field holding a variable number of token vectors per document, for the late-interaction rescore (Issue #1351); see Multi-Vector Fields. dimension is the dimension of each token vector and must be greater than 0; distance is "cosine" (default) or "dot_product". Both are checked when the field is added, which throws with code InvalidArg. storage sets the on-disk element kind of each token vector (Issue #1346): "f32" (default, exact), "f16" (2x smaller), or "int8" (~4x smaller); an unrecognized value also throws InvalidArg. embedder names a token-level ("candle_colbert") embedder, which embeds text values and rescore query text. The field is not searchable on its own, and its token vectors are not stored, so getDocuments and search results do not return them. |
addEmbedder(name, config) | Register a named embedder. See Embedders. |
addAnalyzer(name, tokenizer, charFilters?, tokenFilters?) | Register a custom analyzer definition. tokenizer is required; charFilters/tokenFilters are optional arrays of objects. Each object uses the same { type: "...", ... } shape as the schema TOML/JSON format (see below) — keys stay snake_case, matching that wire format. A name reserved for a built-in analyzer (standard, keyword, english, simple, noop) throws, and so does Index.create with a new index’s schema that defines one (e.g. loaded with fromToml). Semantic validity (e.g. a malformed regex) is checked when the schema is used to build an Index, not here. |
analyzerNames() | Return the names of custom analyzers registered via addAnalyzer or loaded from TOML. |
Schema.fromToml(tomlStr) (static) | Parse a schema from a TOML string, in the same format laurus-cli create index --schema accepts. |
Schema.fromTomlFile(path) (static) | Load a schema from a TOML file. |
toToml() | Serialize this schema to a TOML string in the same format laurus-cli accepts. |
toTomlFile(path) | Write this schema to a TOML file. |
setDefaultFields(fields) | Set default search fields. |
setDynamicFieldPolicy(policy) | Set how undeclared fields are handled. policy is "strict", "dynamic" (default), or "ignore". See notes below. |
dynamicFieldPolicy() | Return the current policy as a lowercase string. |
fieldNames() | Return all field names. |
toString() | Return a string representation of the schema ("Schema(fields=[...])"). |
Vector quantization & rerank storage (HNSW fields):
quantizer—"scalar_8bit"(default, 4× compression) or"product_quantization"for higher compression. Product quantization requiressubvectorCount(must dividedimension).rerankStorage— set to"f32"to write a full-precision*.hnsw.f32sidecar enabling exact Stage-2 rerank; omit to keep the int8-only segment.pqCodebookPath— storage-relative file name of a shared PQ codebook (Issue #631), trained once via thelaurus train pq-codebookCLI command. Only meaningful withquantizer: "product_quantization"; commits then encode against the pre-trained codebook instead of re-training k-means per segment. Omit to keep per-segment training.
Every add*Field method above throws when name starts with _ (other than _id), and adds nothing. A schema loaded with fromToml / fromTomlFile keeps accepting such a field, so a persisted schema still loads; creating a new Index from it throws. See Field Naming.
Dynamic field policy
Controls what happens when a document is ingested with field names that are not declared in the schema:
"strict"— Reject the document."dynamic"(default) — Infer a type for each undeclared field and add it to the schema. Warning: integer fields silently truncate incoming float values (3.14→3). Use"strict"if you need to reject such type mismatches."ignore"— Silently drop the undeclared fields.
See Schema & Fields for the full behaviour matrix.
Analyzer components
Used by addAnalyzer(name, tokenizer, charFilters?, tokenFilters?) and by
the [analyzers.<name>] TOML section. tokenizer is a single object;
charFilters/tokenFilters are arrays of objects, applied in array order.
See Schema Format Reference → Analyzers for the canonical description of each component.
Tokenizers (tokenizer, exactly one):
type | Required keys | Optional keys |
|---|---|---|
"whitespace" | – | – |
"unicode_word" | – | – |
"regex" | – | pattern (default \w+), gaps (default false) |
"ngram" | min_gram, max_gram | – |
"lindera" | mode, dict | user_dict |
"whole" | – | – |
Char filters (charFilters, applied to raw text before tokenization):
type | Required keys | Optional keys |
|---|---|---|
"unicode_normalization" | form ("nfc"/"nfd"/"nfkc"/"nfkd") | – |
"pattern_replace" | pattern, replacement | – |
"mapping" | mapping (object of string replacements) | – |
"japanese_iteration_mark" | – | kanji (default true), kana (default true) |
Token filters (tokenFilters, applied to the token stream after tokenization):
type | Required keys | Optional keys |
|---|---|---|
"lowercase" | – | – |
"stop" | – | words (default: English stop words) |
"stem" | – | stem_type ("porter"/"simple"/"identity") |
"boost" | boost | – |
"limit" | limit | – |
"strip" | – | – |
"remove_empty" | – | – |
"flatten_graph" | – | – |
const schema = new Schema();
schema.addAnalyzer(
"ja_ipadic",
{ type: "lindera", mode: "normal", dict: "/var/lib/lindera/ipadic" },
[
{ type: "unicode_normalization", form: "nfkc" },
{ type: "japanese_iteration_mark" },
],
[{ type: "lowercase" }],
);
schema.addTextField("title", true, true, true, true, "ja_ipadic");
Embedders
Used by addEmbedder(name, config) and by the [embedders.<name>] TOML
section. config is an object whose type key selects the backend; like the
analyzer components, its keys stay snake_case, matching the schema TOML/JSON
format.
See Schema Format Reference → Embedders for the canonical description of each type.
type | Required keys | Optional keys | Feature flag |
|---|---|---|---|
"precomputed" | – | – | (always available) |
"candle_bert" | model | – | embeddings-candle |
"candle_clip" | model | – | embeddings-multimodal |
"openai" | model | – | embeddings-openai |
"candle_colbert" | model | revision, query_maxlen, doc_maxlen | embeddings-candle |
"candle_colbert" runs a BERT-based ColBERT checkpoint (such as
"colbert-ir/colbertv2.0") and produces one vector per token, so only a
multi-vector field (addMultiVectorField) can use it; a multi-vector field in
turn accepts only "candle_colbert" or "precomputed", and any other
combination makes Index.create throw. revision pins a branch, tag or
commit of the model repository (pin a commit, so that re-embedding reproduces
the same vectors); query_maxlen and doc_maxlen override the checkpoint’s
query and document token lengths.
addEmbedder throws embedder config must be an object when config is not
an object, and invalid embedder config: ... when type is missing or
unknown or a required key is missing. A type whose feature flag the binding
was not built with still registers, but Index.create throws.
const schema = new Schema();
schema.addEmbedder("colbert", {
type: "candle_colbert",
model: "answerdotai/answerai-colbert-small-v1",
revision: "934fa8bb4ce2284f4c2baa232d81aca4d076fa5e",
});
schema.addTextField("body");
schema.addMultiVectorField("body_colbert", 96, "cosine", "colbert");
Distance metrics
| Value | Description |
|---|---|
"cosine" | Cosine similarity (default) |
"euclidean" | Euclidean distance |
"dot_product" | Dot product |
"manhattan" | Manhattan distance |
"angular" | Angular distance |
Query classes
TermQuery
new TermQuery(field: string, term: string)
Matches documents containing the exact term in the given field.
PhraseQuery
new PhraseQuery(field: string, terms: string[])
Matches documents containing the terms in order.
FuzzyQuery
new FuzzyQuery(field: string, term: string, maxEdits?: number)
Approximate match allowing up to maxEdits edit-distance
errors (default 2).
WildcardQuery
new WildcardQuery(field: string, pattern: string)
Pattern match. * matches any sequence, ? matches one
character.
NumericRangeQuery
new NumericRangeQuery(
field: string,
min?: number | null,
max?: number | null,
numericType?: "integer" | "float",
)
Matches numeric values in [min, max]. Pass null (or omit) for an
open bound. numericType selects the underlying range type
("integer" (default) or "float"); other values throw.
DateTimeRangeQuery
new DateTimeRangeQuery(
field: string,
min?: string | null,
max?: string | null,
)
Matches DateTime values in [min, max] (both bounds inclusive). Pass
null (or omit) for an open bound. Bounds are string literals in any form
the query DSL accepts: RFC 3339 ("2024-01-01T09:00:00+09:00", normalized
to UTC), naive "YYYY-MM-DDTHH:MM:SS[.fff]" (UTC), or "YYYY-MM-DD"
(midnight UTC); pass date.toISOString() for a Date. A malformed bound
throws an Error at construction. Attach it with
BooleanQuery.mustDateTimeRange / shouldDateTimeRange /
mustNotDateTimeRange and SearchRequest.setLexicalDateTimeRange /
setFilterDateTimeRange.
GeoDistanceQuery
GeoDistanceQuery.withinRadius(
field: string, lat: number, lon: number, distanceM: number,
): GeoDistanceQuery
Geographic distance (radius) search.
GeoBoundingBoxQuery
GeoBoundingBoxQuery.withinBoundingBox(
field: string,
minLat: number, minLon: number,
maxLat: number, maxLon: number,
): GeoBoundingBoxQuery
Geographic bounding-box search.
Geo3dDistanceQuery
Geo3dDistanceQuery.withinSphere(
field: string,
x: number, y: number, z: number,
distanceM: number,
): Geo3dDistanceQuery
Sphere search over a 3D ECEF point field. Returns documents whose (x, y, z)
coordinate is within distanceM metres of the centre. See
Geo3d concepts for ECEF theory.
Geo3dBoundingBoxQuery
Geo3dBoundingBoxQuery.withinBox(
field: string,
minX: number, minY: number, minZ: number,
maxX: number, maxY: number, maxZ: number,
): Geo3dBoundingBoxQuery
Axis-aligned 3D bounding-box search.
Geo3dNearestQuery
Geo3dNearestQuery.kNearest(
field: string,
x: number, y: number, z: number,
k: number,
initialRadiusM?: number,
maxRadiusM?: number,
): Geo3dNearestQuery
k-nearest-neighbour search over a 3D ECEF point field. The optional
initialRadiusM and maxRadiusM parameters tune the iterative-expansion
search cone.
BooleanQuery
class BooleanQuery {
constructor();
// For each query type X in
// { Term, Phrase, Fuzzy, Wildcard, NumericRange, DateTimeRange,
// GeoDistance, GeoBoundingBox,
// Geo3dDistance, Geo3dBoundingBox, Geo3dNearest,
// Boolean, Span }:
mustX(query: X): void;
shouldX(query: X): void;
mustNotX(query: X): void;
}
Compound boolean query with MUST / SHOULD / MUST_NOT clauses. Each clause
takes an instance of a specific query class — for example,
mustTerm(new TermQuery("body", "rust")) or
shouldGeo3dNearest(Geo3dNearestQuery.kNearest(...)).
The Node.js binding exposes 39 per-type methods (13 query types × 3
polarities) instead of a single polymorphic must(query) because of a
limitation in napi-derive’s validation of Either<&T, ...> arguments
for classes with js_name overrides.
must clauses all have to match; mustNot clauses must not match.
should clauses contribute to scoring; at least one of them must match if
there are no must clauses.
const bq = new BooleanQuery();
bq.mustTerm(new TermQuery("body", "programming"));
bq.mustNotTerm(new TermQuery("title", "python"));
bq.shouldFuzzy(new FuzzyQuery("body", "data", 1));
SpanQuery
SpanQuery.term(field: string, term: string): SpanQuery
SpanQuery.near(
field: string, terms: string[],
slop?: number, ordered?: boolean,
): SpanQuery
SpanQuery.nearSpans(
field: string, clauses: SpanQuery[],
slop?: number, ordered?: boolean,
): SpanQuery
SpanQuery.containing(
field: string, big: SpanQuery, little: SpanQuery,
): SpanQuery
SpanQuery.within(
field: string,
include: SpanQuery, exclude: SpanQuery, distance: number,
): SpanQuery
Positional/proximity span queries.
VectorQuery
new VectorQuery(field: string, vector: number[])
Nearest-neighbor search using a pre-computed embedding vector.
VectorTextQuery
new VectorTextQuery(field: string, text: string)
Converts text to an embedding at query time. Requires an
embedder configured on the index.
SearchRequest
Full-featured search request for advanced control.
interface SearchRequestOptions {
queryDsl?: string;
limit?: number; // default 10
offset?: number; // default 0
highlight?: HighlightOptions;
rescore?: RescoreOptions;
}
class SearchRequest {
constructor(options?: SearchRequestOptions);
}
Construct with primitive options first; attach polymorphic clauses with
the per-type setters below. As with BooleanQuery, the binding exposes
per-type setters because of napi-derive’s limitation on Either<&T, ...>
arguments. highlight and rescore are plain data (not class-instance
unions), so they live directly on SearchRequestOptions rather than behind a
setter.
DSL and fusion setters
| Method | Description |
|---|---|
setQueryDsl(dsl: string) | Set a DSL string query. |
setRrfFusion(rrf: RRF) | Use RRF fusion. |
setWeightedSumFusion(ws: WeightedSum) | Use weighted-sum fusion. |
Vector setters
| Method | Description |
|---|---|
setVectorQuery(query: VectorQuery) | Set a pre-computed vector query. |
setVectorTextQuery(query: VectorTextQuery) | Set a text-based vector query (auto-embedded by the configured embedder). |
Lexical setters (per type)
For each query type X in { Term, Phrase, Fuzzy, Wildcard, NumericRange, DateTimeRange, GeoDistance, GeoBoundingBox, Geo3dDistance, Geo3dBoundingBox, Geo3dNearest, Boolean, Span }, the request exposes:
| Method | Description |
|---|---|
setLexicalX(query: X) | Set the lexical component for an explicit hybrid request. |
setFilterX(query: X) | Set the post-scoring filter component. |
That is, 26 per-type setters in total (13 lexical + 13 filter), in addition
to the DSL, vector, and fusion setters above — for example
setLexicalNumericRange(q) / setFilterNumericRange(q) and
setLexicalDateTimeRange(q) / setFilterDateTimeRange(q).
const req = new SearchRequest({ limit: 5 });
req.setLexicalTerm(new TermQuery("title", "rust"));
req.setVectorQuery(new VectorQuery("embedding", [0.1, 0.2, 0.3, 0.4]));
req.setRrfFusion(new RRF(60.0));
const results = await index.searchWithRequest(req);
Highlighting
search, searchTerm, searchBatch and SearchRequestOptions accept an optional highlight object (Issue #1134):
interface HighlightOptions {
fields: string[];
fragmentSize?: number; // default 150
maxFragments?: number; // default 5
tag?: string; // default "mark"
cssClass?: string;
requireFieldMatch?: boolean; // default true
}
Only fields is required. Highlighting follows the query passed to the same call, and only stored: true text fields can be highlighted — a field that isn’t stored, isn’t a text field, or had no match is simply absent from the result’s highlights object.
const results = await index.search("body:rust", 10, 0, { fields: ["body"], tag: "em" });
// results[0].highlights => { body: ["<em>Rust</em> is a systems programming language."] }
Late-interaction rescore
search (its trailing rescore argument) and SearchRequestOptions accept
an optional rescore object (Issue #1351). It reorders the top results of the
first stage — lexical, vector or hybrid — by ColBERT-style late interaction
(MaxSim) against a
multi-vector field.
See Vector Search → Late-Interaction Rescore
for how it works.
// Exported from index.d.ts as JsRescoreOptions.
interface RescoreOptions {
field: string; // a multi-vector field
vectors?: number[][]; // the query's token vectors
text?: string; // query text, embedded by the field's embedder
windowSize?: number; // default 100, at most 10,000
}
| Field | Type | Default | Description |
|---|---|---|---|
field | string | – | The multi-vector field to score against. |
vectors | number[][] | – | The query’s token vectors: 1 to 1,024 of them, each of the field’s dimension, with finite values. Compute them with the model that produced the documents’ token vectors. |
text | string | – | Query text, embedded as a query by the field’s token-level ("candle_colbert") embedder. Must not be blank. |
windowSize | number | 100 | How many top first-stage results to rescore, 1 to 10,000. |
Set exactly one of vectors and text; otherwise the search rejects with
code InvalidArg and the message
rescore needs exactly one of vectors or text. The engine checks the other
values when the search runs, and rejects with code InvalidArg and a message
containing rescore: ... when field is unknown or not a multi-vector field,
a text query is blank or the field has no token-level embedder, the query
holds too few or too many vectors or one of the wrong dimension, or
windowSize is out of range.
Ordering and scores:
- The top
windowSizefirst-stage results are sorted by MaxSim, highest first, and a rescored result’sscoreis its MaxSim. - Window results without token vectors in the field follow in their first-stage order, then the results beyond the window, unchanged. Both keep their first-stage score, which is not comparable with a MaxSim.
searchTerm, searchVector, searchVectorText and searchBatch take no
rescore argument; build a SearchRequest instead, which rescores any first
stage, including a hybrid one.
// A DSL search, rescored with the query's token vectors.
const results = await index.search("title:rust", 10, 0, undefined, {
field: "tokens",
vectors: [[1, 0], [0, 1]],
});
// A hybrid search, rescored with query text (the field needs a
// "candle_colbert" embedder).
const req = new SearchRequest({
limit: 10,
rescore: { field: "body_colbert", text: "how do lifetimes work", windowSize: 50 },
});
req.setLexicalTerm(new TermQuery("body", "lifetimes"));
req.setVectorQuery(new VectorQuery("body_vec", queryEmbedding));
req.setRrfFusion(new RRF(60.0));
const reranked = await index.searchWithRequest(req);
SearchResult
Returned by search methods as an array.
interface SearchResult {
id: string; // External document identifier
score: number; // Relevance score
document: object | null; // Retrieved fields, or null if not stored
highlights: Record<string, string[]>; // Highlighted fragments per requested field
}
highlights maps each field named in highlight.fields to its highlighted fragments (best first); a field that did not highlight is absent from the object, and highlights is {} when highlight was not requested. See Highlighting.
Fusion algorithms
RRF
new RRF(k?: number) // default 60.0
Reciprocal Rank Fusion. Merges lexical and vector result lists by rank position.
WeightedSum
new WeightedSum(
lexicalWeight?: number, // default 0.5
vectorWeight?: number, // default 0.5
)
Normalises both score lists independently, then combines them.
Text analysis
SynonymDictionary
class SynonymDictionary {
constructor();
addSynonymGroup(terms: string[]): void;
}
WhitespaceTokenizer
class WhitespaceTokenizer {
constructor();
tokenize(text: string): Token[];
}
SynonymGraphFilter
class SynonymGraphFilter {
constructor(
dictionary: SynonymDictionary,
keepOriginal?: boolean, // default true
boost?: number, // default 1.0
);
apply(tokens: Token[]): Token[];
}
Token
interface Token {
text: string;
position: number;
startOffset: number;
endOffset: number;
boost: number;
stopped: boolean;
positionIncrement: number;
positionLength: number;
tokenType?: string;
}
startOffset and endOffset are UTF-8 byte offsets into the original text. For non-ASCII text they are not JavaScript string (UTF-16) indices; slice the encoded text instead: new TextDecoder().decode(new TextEncoder().encode(text).slice(tok.startOffset, tok.endOffset)).
tokenType is one of "alphanum", "num", "cjk", "katakana", "hiragana", "hangul", "punctuation", "whitespace", "synonym", "email", "url" and "other". SynonymGraphFilter.apply keeps each token’s offsets and type, and gives each synonym it inserts the type "synonym" and the offsets of the words it replaces. It throws on any other tokenType.
A token built by hand may leave tokenType out. A multi-word synonym then matches only if the words’ offsets touch, so set tokenType: "alphanum" on words separated by spaces.
Field value types
JavaScript values are automatically converted to Laurus
DataValue types:
| JavaScript type | Laurus type | Notes |
|---|---|---|
null | Null | |
boolean | Bool | |
number (integer) | Int64 | |
number (float) | Float64 | |
string | Text | ISO 8601 strings become DateTime |
number[] (all integers) | Int64Array | Multi-valued integer field; vector fields cast the array to f32. An empty array is an empty Int64Array |
number[] | Float64Array | Multi-valued float field (integers widened); vector fields cast the array to f32 |
{ lat, lon } | Geo | Two number values |
{ x, y, z } | GeoEcef | Three number values, meters (3D ECEF Cartesian) |
{ lat, lon }[] | GeoArray | Array of { lat, lon } objects; requires multiValued: true on the field |
{ x, y, z }[] | GeoEcefArray | Array of { x, y, z } objects; requires multiValued: true on the field |
number[][] | VectorArray | Token vectors of a multi-vector field (addMultiVectorField): 1 to 8,192 arrays of the field’s dimension. All inner arrays must have the same length, otherwise ingestion rejects with token vectors must share one dimension. A field with a token-level embedder also accepts a string, which is embedded into token vectors. Not stored, so absent from getDocuments and search results |
string[] (all RFC 3339) | DateTimeArray | Array of RFC 3339 datetime strings (same infer_from_json rules as the HTTP gateway); requires multiValued: true on the field |
string[] (not all RFC 3339) | TextArray | Array of strings (Issue #1175); requires multiValued: true on the field. Read back as string[]. On a declared multi-valued Bytes field, the same array of base64 strings is instead decoded element-wise into BytesArray by the schema-aware coercion (Issue #1176) |
boolean[] | BoolArray | Array of booleans (same infer_from_json rules as the HTTP gateway); requires multiValued: true on the field. A mixed array such as [true, 1] is rejected |