Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

API Reference

Index

The primary entry point. Wraps the Laurus search engine.

class Index {
  static create(
    path?: string | null,
    schema?: Schema,
    walSyncPolicy?: WalSyncPolicy,
    commitPolicy?: CommitPolicy,
  ): Promise<Index>;
}

Factory method

ParameterTypeDefaultDescription
pathstring | nullnullDirectory for persistent storage. null creates an in-memory index. When given, the directory follows the <path>/schema.toml + <path>/store/ layout laurus-cli create index/--index-dir uses, so an index built here can also be opened by the CLI (and vice versa). See below.
schemaSchemaemptySchema definition. Only meaningful when creating a new file-backed index (or an in-memory one); must be omitted when reopening an existing file-backed index – the persisted schema is loaded instead. An empty schema is used when omitted for a new index.
walSyncPolicyWalSyncPolicyper-recordWAL durability policy. Omit to keep the default per-record fsync. See WAL sync policy / durability.
commitPolicyCommitPolicymanualAuto-commit policy. Omit to keep the default manual policy (the caller drives every commit()). See Commit policy / auto-commit.

Creating vs. reopening a file-backed index (path given): if <path>/schema.toml does not yet exist, this call creates a new index and persists schema (or an empty schema, if omitted) to it. If <path>/schema.toml already exists, this call reopens the index – schema must be omitted, or the call throws (it would be ambiguous which schema should win). It also throws if path contains an index in the layout used before this convention was introduced (segment files directly under path, no schema.toml).

Methods

MethodDescription
putDocument(id, doc)Upsert a document. Replaces all existing versions.
addDocument(id, doc)Append a document chunk without removing existing versions.
putDocuments(docs)Batched upsert. docs is an Array<[id, doc]> applied in order with one WAL fsync per batch (duplicate ids dedup, last wins). Fails fast at the first bad entry; the applied prefix is not rolled back.
addDocuments(docs)Batched chunk append. Like putDocuments but repeated ids accumulate as separate versions.
getDocuments(id)Return all stored versions for the given ID.
deleteDocuments(id)Delete all versions for the given ID.
commit()Flush writes and make pending changes searchable.
flushWal()Force a durable WAL barrier. See WAL sync policy / durability.
search(query, limit?, offset?, highlight?, rescore?)Search with a DSL string. rescore reorders the top results by late interaction; see Late-interaction rescore.
searchTerm(field, term, limit?, offset?, highlight?)Search with an exact term match.
searchVector(field, vector, limit?, offset?)Search with a pre-computed vector.
searchVectorText(field, text, limit?, offset?)Search with text (auto-embedded).
searchWithRequest(request)Search with a SearchRequest.
searchBatch(queries, limit?, offset?, highlight?)Execute multiple DSL string queries in parallel. results[i] corresponds to queries[i]. Returns Promise<Array<Array<JsSearchResult>>>. Empty input returns []. highlight applies identically to every query.
stats()Return index statistics (documentCount, vectorFields).

All document methods and search methods are async and return Promises. stats() is synchronous.

stats() returns an object shaped like:

interface IndexStats {
  documentCount: number;
  vectorFields: Record<string, { count: number; dimension: number }>;
}

WAL sync policy / durability

For a persistent index, every write is appended to a write-ahead log (WAL). By default the WAL is fsync-ed on every record, so each write is fully durable as soon as its Promise resolves. Index.create accepts an optional walSyncPolicy to trade some durability for higher write throughput, and flushWal() forces a durable barrier on demand.

class WalSyncPolicy {
  static perRecord(): WalSyncPolicy;
  static group(
    maxRecords?: number,
    maxBytes?: number,
    maxIntervalMs?: number,
  ): WalSyncPolicy;
}
ConstructorDescription
WalSyncPolicy.perRecord()Default. fsync after every WAL record; fully durable per write.
WalSyncPolicy.group(...)Group commit. Batches fsync across writes.

group(...) parameters (omit any argument to keep its default):

ParameterDefaultDescription
maxRecords1024Flush once this many records have accumulated.
maxBytes1048576 (1 MiB)Flush once this many unsynced bytes have accumulated.
maxIntervalMsnoneOptional periodic flush timer (milliseconds). Omit to disable the timer.

With group commit the WAL is flushed when either maxRecords or maxBytes is reached, and always at commit(). A crash can lose up to the last unsynced batch — the same trade-off as SQLite’s synchronous = NORMAL. Call flushWal() to force everything written so far to disk without a full commit().

MethodDescription
flushWal()Force a durable WAL barrier now. Returns Promise<void>.
import { Index, WalSyncPolicy } from "laurus-nodejs";

// Opt into group commit with a 1-second periodic flush timer.
const policy = WalSyncPolicy.group(4096, undefined, 1000);
const index = await Index.create("./myindex", schema, policy);

for (let i = 0; i < 10000; i++) {
  await index.putDocument(`doc${i}`, { title: `Document ${i}` });
}

// Force a durable barrier without committing yet.
await index.flushWal();

await index.commit(); // also flushes the WAL

Omit walSyncPolicy (or pass WalSyncPolicy.perRecord()) to keep the default, fully durable behaviour.

Commit policy / auto-commit

By default the caller drives every commit(), so pending changes become searchable only when you call it explicitly. Index.create accepts an optional commitPolicy to let the engine auto-commit on an ingestion-driven cadence instead, so writes are materialized without an explicit commit().

class CommitPolicy {
  static manual(): CommitPolicy;
  static everyDocs(n: number): CommitPolicy;
  static intervalMs(ms: number): CommitPolicy;
}
ConstructorDescription
CommitPolicy.manual()Default. No auto-commit; the caller drives every commit().
CommitPolicy.everyDocs(n)Auto-commit after every n applied documents.
CommitPolicy.intervalMs(ms)Auto-commit at least every ms milliseconds via a background timer (default: none). Native-only; a no-op on wasm (no background threads).

With everyDocs(n) the engine commits after every n applied documents, counted across both singular and batch ingest — including every n documents within a single batch. Passing everyDocs(0) is valid and disables auto-commit, which is equivalent to CommitPolicy.manual().

intervalMs(ms) is the time-based counterpart of everyDocs(n): a background timer auto-commits at least every ms milliseconds, so a trailing partial batch is materialized even while ingestion is idle. It is native-only — on wasm there are no background threads, so the engine treats intervalMs as a no-op (the factory still constructs the value, but no timed commit happens under WebAssembly).

commitPolicy is orthogonal to walSyncPolicy: walSyncPolicy governs WAL fsync durability, while commitPolicy governs when the stores materialize pending changes into a searchable commit. You can combine them freely.

import { Index, CommitPolicy } from "laurus-nodejs";

// Auto-commit after every 1000 applied documents.
const index = await Index.create(
  null,
  schema,
  undefined,
  CommitPolicy.everyDocs(1000),
);

for (let i = 0; i < 10000; i++) {
  await index.putDocument(`doc${i}`, { title: `Document ${i}` });
}
// The engine has already committed ten times; no explicit commit() needed.

Omit commitPolicy (or pass CommitPolicy.manual()) to keep the default, caller-driven commit behaviour.


Schema

Defines the fields and index types for an Index.

class Schema {
  constructor();
}

Field methods

MethodDescription
addTextField(name, stored?, indexed?, termVectors?, docValues?, analyzer?, multiValued?, positionIncrementGap?)Full-text field (inverted index, BM25). docValues controls whether the value is also copied into DocValues, the column-oriented store sort/facet/aggregation read from (Issue #1047, default true); takes effect only when stored is also true. Pass multiValued: true to accept an array of strings (Issue #1175): a term query matches if any element contains the term, a phrase query never spans two elements unless its slop reaches positionIncrementGap (default 100; 0 numbers the elements as if concatenated), and values are read back as an array of strings. analyzer is the name of a parameter-less built-in ("standard", "english", "keyword", "simple", "noop") or any custom name registered via addAnalyzer. For the parameterised Japanese preset (which requires a Lindera dictionary path), register a custom analyzer with a lindera tokenizer and reference it by name.
addIntegerField(name, stored?, indexed?, multiValued?, docValues?)64-bit integer field. Pass multiValued: true to accept arrays of integers (range queries match if any value satisfies the predicate). See docValues above.
addFloatField(name, stored?, indexed?, multiValued?, docValues?)64-bit float field. Pass multiValued: true to accept arrays of floats (range queries match if any value satisfies the predicate). See docValues above.
addBooleanField(name, stored?, indexed?, multiValued?, docValues?)Boolean field. Pass multiValued: true to accept an array of booleans (a term query such as flags:true matches if any element equals the value; values are read back as an array of booleans). See docValues above.
addBytesField(name, stored?, multiValued?)Raw bytes field. No docValues option: a Bytes value is never written to DocValues regardless. Pass multiValued: true to accept an array of base64 strings (Issue #1176), decoded element-wise by the schema-aware coercion the same way a single base64 string is; Bytes is never indexed, so unlike every other multiValued option this has no query-matching semantics — it only governs the stored shape and ingestion arity. Values are read back as an array of byte-integer arrays with MIME dropped, the same input/output asymmetry the scalar field already has.
addGeoField(name, stored?, indexed?, multiValued?, docValues?)Geographic coordinate field. Pass multiValued: true to accept an array of { lat, lon } objects (distance / bounding-box queries match if any point satisfies the predicate). See docValues above.
addGeo3dField(name, stored?, indexed?, multiValued?, docValues?)3D ECEF Cartesian point field (x, y, z in metres). Pass multiValued: true to accept an array of { x, y, z } objects (distance / bounding-box / nearest queries match if any point satisfies the predicate). See Geo3d concepts and docValues above.
addDatetimeField(name, stored?, indexed?, multiValued?, docValues?)UTC datetime field. Pass multiValued: true to accept an array of RFC 3339 strings (range queries match if any instant satisfies the predicate; values are read back as an array of RFC 3339 strings normalized to UTC). See docValues above.
addHnswField(name, dimension, distance?, m?, efConstruction?, defaultEfSearch?, embedder?, quantizer?, subvectorCount?, rerankStorage?, pqCodebookPath?, baseWeight?)HNSW vector field. baseWeight sets this field’s relative scoring priority when searched alongside other vector fields (Issue #1084); see Vector Search → Weights.
addFlatField(name, dimension, distance?, embedder?, baseWeight?)Flat (brute-force) vector field.
addIvfField(name, dimension, distance?, nClusters?, nProbe?, embedder?, baseWeight?)IVF vector field.
addMultiVectorField(name, dimension, distance?, embedder?, storage?)Multi-vector field holding a variable number of token vectors per document, for the late-interaction rescore (Issue #1351); see Multi-Vector Fields. dimension is the dimension of each token vector and must be greater than 0; distance is "cosine" (default) or "dot_product". Both are checked when the field is added, which throws with code InvalidArg. storage sets the on-disk element kind of each token vector (Issue #1346): "f32" (default, exact), "f16" (2x smaller), or "int8" (~4x smaller); an unrecognized value also throws InvalidArg. embedder names a token-level ("candle_colbert") embedder, which embeds text values and rescore query text. The field is not searchable on its own, and its token vectors are not stored, so getDocuments and search results do not return them.
addEmbedder(name, config)Register a named embedder. See Embedders.
addAnalyzer(name, tokenizer, charFilters?, tokenFilters?)Register a custom analyzer definition. tokenizer is required; charFilters/tokenFilters are optional arrays of objects. Each object uses the same { type: "...", ... } shape as the schema TOML/JSON format (see below) — keys stay snake_case, matching that wire format. A name reserved for a built-in analyzer (standard, keyword, english, simple, noop) throws, and so does Index.create with a new index’s schema that defines one (e.g. loaded with fromToml). Semantic validity (e.g. a malformed regex) is checked when the schema is used to build an Index, not here.
analyzerNames()Return the names of custom analyzers registered via addAnalyzer or loaded from TOML.
Schema.fromToml(tomlStr) (static)Parse a schema from a TOML string, in the same format laurus-cli create index --schema accepts.
Schema.fromTomlFile(path) (static)Load a schema from a TOML file.
toToml()Serialize this schema to a TOML string in the same format laurus-cli accepts.
toTomlFile(path)Write this schema to a TOML file.
setDefaultFields(fields)Set default search fields.
setDynamicFieldPolicy(policy)Set how undeclared fields are handled. policy is "strict", "dynamic" (default), or "ignore". See notes below.
dynamicFieldPolicy()Return the current policy as a lowercase string.
fieldNames()Return all field names.
toString()Return a string representation of the schema ("Schema(fields=[...])").

Vector quantization & rerank storage (HNSW fields):

  • quantizer — "scalar_8bit" (default, 4× compression) or "product_quantization" for higher compression. Product quantization requires subvectorCount (must divide dimension).
  • rerankStorage — set to "f32" to write a full-precision *.hnsw.f32 sidecar enabling exact Stage-2 rerank; omit to keep the int8-only segment.
  • pqCodebookPath — storage-relative file name of a shared PQ codebook (Issue #631), trained once via the laurus train pq-codebook CLI command. Only meaningful with quantizer: "product_quantization"; commits then encode against the pre-trained codebook instead of re-training k-means per segment. Omit to keep per-segment training.

Every add*Field method above throws when name starts with _ (other than _id), and adds nothing. A schema loaded with fromToml / fromTomlFile keeps accepting such a field, so a persisted schema still loads; creating a new Index from it throws. See Field Naming.

Dynamic field policy

Controls what happens when a document is ingested with field names that are not declared in the schema:

  • "strict" — Reject the document.
  • "dynamic" (default) — Infer a type for each undeclared field and add it to the schema. Warning: integer fields silently truncate incoming float values (3.14 → 3). Use "strict" if you need to reject such type mismatches.
  • "ignore" — Silently drop the undeclared fields.

See Schema & Fields for the full behaviour matrix.

Analyzer components

Used by addAnalyzer(name, tokenizer, charFilters?, tokenFilters?) and by the [analyzers.<name>] TOML section. tokenizer is a single object; charFilters/tokenFilters are arrays of objects, applied in array order.

See Schema Format Reference → Analyzers for the canonical description of each component.

Tokenizers (tokenizer, exactly one):

typeRequired keysOptional keys
"whitespace"––
"unicode_word"––
"regex"–pattern (default \w+), gaps (default false)
"ngram"min_gram, max_gram–
"lindera"mode, dictuser_dict
"whole"––

Char filters (charFilters, applied to raw text before tokenization):

typeRequired keysOptional keys
"unicode_normalization"form ("nfc"/"nfd"/"nfkc"/"nfkd")–
"pattern_replace"pattern, replacement–
"mapping"mapping (object of string replacements)–
"japanese_iteration_mark"–kanji (default true), kana (default true)

Token filters (tokenFilters, applied to the token stream after tokenization):

typeRequired keysOptional keys
"lowercase"––
"stop"–words (default: English stop words)
"stem"–stem_type ("porter"/"simple"/"identity")
"boost"boost–
"limit"limit–
"strip"––
"remove_empty"––
"flatten_graph"––
const schema = new Schema();
schema.addAnalyzer(
  "ja_ipadic",
  { type: "lindera", mode: "normal", dict: "/var/lib/lindera/ipadic" },
  [
    { type: "unicode_normalization", form: "nfkc" },
    { type: "japanese_iteration_mark" },
  ],
  [{ type: "lowercase" }],
);
schema.addTextField("title", true, true, true, true, "ja_ipadic");

Embedders

Used by addEmbedder(name, config) and by the [embedders.<name>] TOML section. config is an object whose type key selects the backend; like the analyzer components, its keys stay snake_case, matching the schema TOML/JSON format.

See Schema Format Reference → Embedders for the canonical description of each type.

typeRequired keysOptional keysFeature flag
"precomputed"––(always available)
"candle_bert"model–embeddings-candle
"candle_clip"model–embeddings-multimodal
"openai"model–embeddings-openai
"candle_colbert"modelrevision, query_maxlen, doc_maxlenembeddings-candle

"candle_colbert" runs a BERT-based ColBERT checkpoint (such as "colbert-ir/colbertv2.0") and produces one vector per token, so only a multi-vector field (addMultiVectorField) can use it; a multi-vector field in turn accepts only "candle_colbert" or "precomputed", and any other combination makes Index.create throw. revision pins a branch, tag or commit of the model repository (pin a commit, so that re-embedding reproduces the same vectors); query_maxlen and doc_maxlen override the checkpoint’s query and document token lengths.

addEmbedder throws embedder config must be an object when config is not an object, and invalid embedder config: ... when type is missing or unknown or a required key is missing. A type whose feature flag the binding was not built with still registers, but Index.create throws.

const schema = new Schema();
schema.addEmbedder("colbert", {
  type: "candle_colbert",
  model: "answerdotai/answerai-colbert-small-v1",
  revision: "934fa8bb4ce2284f4c2baa232d81aca4d076fa5e",
});
schema.addTextField("body");
schema.addMultiVectorField("body_colbert", 96, "cosine", "colbert");

Distance metrics

ValueDescription
"cosine"Cosine similarity (default)
"euclidean"Euclidean distance
"dot_product"Dot product
"manhattan"Manhattan distance
"angular"Angular distance

Query classes

TermQuery

new TermQuery(field: string, term: string)

Matches documents containing the exact term in the given field.

PhraseQuery

new PhraseQuery(field: string, terms: string[])

Matches documents containing the terms in order.

FuzzyQuery

new FuzzyQuery(field: string, term: string, maxEdits?: number)

Approximate match allowing up to maxEdits edit-distance errors (default 2).

WildcardQuery

new WildcardQuery(field: string, pattern: string)

Pattern match. * matches any sequence, ? matches one character.

NumericRangeQuery

new NumericRangeQuery(
  field: string,
  min?: number | null,
  max?: number | null,
  numericType?: "integer" | "float",
)

Matches numeric values in [min, max]. Pass null (or omit) for an open bound. numericType selects the underlying range type ("integer" (default) or "float"); other values throw.

DateTimeRangeQuery

new DateTimeRangeQuery(
  field: string,
  min?: string | null,
  max?: string | null,
)

Matches DateTime values in [min, max] (both bounds inclusive). Pass null (or omit) for an open bound. Bounds are string literals in any form the query DSL accepts: RFC 3339 ("2024-01-01T09:00:00+09:00", normalized to UTC), naive "YYYY-MM-DDTHH:MM:SS[.fff]" (UTC), or "YYYY-MM-DD" (midnight UTC); pass date.toISOString() for a Date. A malformed bound throws an Error at construction. Attach it with BooleanQuery.mustDateTimeRange / shouldDateTimeRange / mustNotDateTimeRange and SearchRequest.setLexicalDateTimeRange / setFilterDateTimeRange.

GeoDistanceQuery

GeoDistanceQuery.withinRadius(
  field: string, lat: number, lon: number, distanceM: number,
): GeoDistanceQuery

Geographic distance (radius) search.

GeoBoundingBoxQuery

GeoBoundingBoxQuery.withinBoundingBox(
  field: string,
  minLat: number, minLon: number,
  maxLat: number, maxLon: number,
): GeoBoundingBoxQuery

Geographic bounding-box search.

Geo3dDistanceQuery

Geo3dDistanceQuery.withinSphere(
  field: string,
  x: number, y: number, z: number,
  distanceM: number,
): Geo3dDistanceQuery

Sphere search over a 3D ECEF point field. Returns documents whose (x, y, z) coordinate is within distanceM metres of the centre. See Geo3d concepts for ECEF theory.

Geo3dBoundingBoxQuery

Geo3dBoundingBoxQuery.withinBox(
  field: string,
  minX: number, minY: number, minZ: number,
  maxX: number, maxY: number, maxZ: number,
): Geo3dBoundingBoxQuery

Axis-aligned 3D bounding-box search.

Geo3dNearestQuery

Geo3dNearestQuery.kNearest(
  field: string,
  x: number, y: number, z: number,
  k: number,
  initialRadiusM?: number,
  maxRadiusM?: number,
): Geo3dNearestQuery

k-nearest-neighbour search over a 3D ECEF point field. The optional initialRadiusM and maxRadiusM parameters tune the iterative-expansion search cone.

BooleanQuery

class BooleanQuery {
  constructor();
  // For each query type X in
  //   { Term, Phrase, Fuzzy, Wildcard, NumericRange, DateTimeRange,
  //     GeoDistance, GeoBoundingBox,
  //     Geo3dDistance, Geo3dBoundingBox, Geo3dNearest,
  //     Boolean, Span }:
  mustX(query: X): void;
  shouldX(query: X): void;
  mustNotX(query: X): void;
}

Compound boolean query with MUST / SHOULD / MUST_NOT clauses. Each clause takes an instance of a specific query class — for example, mustTerm(new TermQuery("body", "rust")) or shouldGeo3dNearest(Geo3dNearestQuery.kNearest(...)).

The Node.js binding exposes 39 per-type methods (13 query types × 3 polarities) instead of a single polymorphic must(query) because of a limitation in napi-derive’s validation of Either<&T, ...> arguments for classes with js_name overrides.

must clauses all have to match; mustNot clauses must not match. should clauses contribute to scoring; at least one of them must match if there are no must clauses.

const bq = new BooleanQuery();
bq.mustTerm(new TermQuery("body", "programming"));
bq.mustNotTerm(new TermQuery("title", "python"));
bq.shouldFuzzy(new FuzzyQuery("body", "data", 1));

SpanQuery

SpanQuery.term(field: string, term: string): SpanQuery
SpanQuery.near(
  field: string, terms: string[],
  slop?: number, ordered?: boolean,
): SpanQuery
SpanQuery.nearSpans(
  field: string, clauses: SpanQuery[],
  slop?: number, ordered?: boolean,
): SpanQuery
SpanQuery.containing(
  field: string, big: SpanQuery, little: SpanQuery,
): SpanQuery
SpanQuery.within(
  field: string,
  include: SpanQuery, exclude: SpanQuery, distance: number,
): SpanQuery

Positional/proximity span queries.

VectorQuery

new VectorQuery(field: string, vector: number[])

Nearest-neighbor search using a pre-computed embedding vector.

VectorTextQuery

new VectorTextQuery(field: string, text: string)

Converts text to an embedding at query time. Requires an embedder configured on the index.


SearchRequest

Full-featured search request for advanced control.

interface SearchRequestOptions {
  queryDsl?: string;
  limit?: number;   // default 10
  offset?: number;  // default 0
  highlight?: HighlightOptions;
  rescore?: RescoreOptions;
}

class SearchRequest {
  constructor(options?: SearchRequestOptions);
}

Construct with primitive options first; attach polymorphic clauses with the per-type setters below. As with BooleanQuery, the binding exposes per-type setters because of napi-derive’s limitation on Either<&T, ...> arguments. highlight and rescore are plain data (not class-instance unions), so they live directly on SearchRequestOptions rather than behind a setter.

DSL and fusion setters

MethodDescription
setQueryDsl(dsl: string)Set a DSL string query.
setRrfFusion(rrf: RRF)Use RRF fusion.
setWeightedSumFusion(ws: WeightedSum)Use weighted-sum fusion.

Vector setters

MethodDescription
setVectorQuery(query: VectorQuery)Set a pre-computed vector query.
setVectorTextQuery(query: VectorTextQuery)Set a text-based vector query (auto-embedded by the configured embedder).

Lexical setters (per type)

For each query type X in { Term, Phrase, Fuzzy, Wildcard, NumericRange, DateTimeRange, GeoDistance, GeoBoundingBox, Geo3dDistance, Geo3dBoundingBox, Geo3dNearest, Boolean, Span }, the request exposes:

MethodDescription
setLexicalX(query: X)Set the lexical component for an explicit hybrid request.
setFilterX(query: X)Set the post-scoring filter component.

That is, 26 per-type setters in total (13 lexical + 13 filter), in addition to the DSL, vector, and fusion setters above — for example setLexicalNumericRange(q) / setFilterNumericRange(q) and setLexicalDateTimeRange(q) / setFilterDateTimeRange(q).

const req = new SearchRequest({ limit: 5 });
req.setLexicalTerm(new TermQuery("title", "rust"));
req.setVectorQuery(new VectorQuery("embedding", [0.1, 0.2, 0.3, 0.4]));
req.setRrfFusion(new RRF(60.0));
const results = await index.searchWithRequest(req);

Highlighting

search, searchTerm, searchBatch and SearchRequestOptions accept an optional highlight object (Issue #1134):

interface HighlightOptions {
  fields: string[];
  fragmentSize?: number;    // default 150
  maxFragments?: number;    // default 5
  tag?: string;             // default "mark"
  cssClass?: string;
  requireFieldMatch?: boolean; // default true
}

Only fields is required. Highlighting follows the query passed to the same call, and only stored: true text fields can be highlighted — a field that isn’t stored, isn’t a text field, or had no match is simply absent from the result’s highlights object.

const results = await index.search("body:rust", 10, 0, { fields: ["body"], tag: "em" });
// results[0].highlights => { body: ["<em>Rust</em> is a systems programming language."] }

Late-interaction rescore

search (its trailing rescore argument) and SearchRequestOptions accept an optional rescore object (Issue #1351). It reorders the top results of the first stage — lexical, vector or hybrid — by ColBERT-style late interaction (MaxSim) against a multi-vector field. See Vector Search → Late-Interaction Rescore for how it works.

// Exported from index.d.ts as JsRescoreOptions.
interface RescoreOptions {
  field: string;          // a multi-vector field
  vectors?: number[][];   // the query's token vectors
  text?: string;          // query text, embedded by the field's embedder
  windowSize?: number;    // default 100, at most 10,000
}
FieldTypeDefaultDescription
fieldstring–The multi-vector field to score against.
vectorsnumber[][]–The query’s token vectors: 1 to 1,024 of them, each of the field’s dimension, with finite values. Compute them with the model that produced the documents’ token vectors.
textstring–Query text, embedded as a query by the field’s token-level ("candle_colbert") embedder. Must not be blank.
windowSizenumber100How many top first-stage results to rescore, 1 to 10,000.

Set exactly one of vectors and text; otherwise the search rejects with code InvalidArg and the message rescore needs exactly one of vectors or text. The engine checks the other values when the search runs, and rejects with code InvalidArg and a message containing rescore: ... when field is unknown or not a multi-vector field, a text query is blank or the field has no token-level embedder, the query holds too few or too many vectors or one of the wrong dimension, or windowSize is out of range.

Ordering and scores:

  • The top windowSize first-stage results are sorted by MaxSim, highest first, and a rescored result’s score is its MaxSim.
  • Window results without token vectors in the field follow in their first-stage order, then the results beyond the window, unchanged. Both keep their first-stage score, which is not comparable with a MaxSim.

searchTerm, searchVector, searchVectorText and searchBatch take no rescore argument; build a SearchRequest instead, which rescores any first stage, including a hybrid one.

// A DSL search, rescored with the query's token vectors.
const results = await index.search("title:rust", 10, 0, undefined, {
  field: "tokens",
  vectors: [[1, 0], [0, 1]],
});

// A hybrid search, rescored with query text (the field needs a
// "candle_colbert" embedder).
const req = new SearchRequest({
  limit: 10,
  rescore: { field: "body_colbert", text: "how do lifetimes work", windowSize: 50 },
});
req.setLexicalTerm(new TermQuery("body", "lifetimes"));
req.setVectorQuery(new VectorQuery("body_vec", queryEmbedding));
req.setRrfFusion(new RRF(60.0));
const reranked = await index.searchWithRequest(req);

SearchResult

Returned by search methods as an array.

interface SearchResult {
  id: string;        // External document identifier
  score: number;     // Relevance score
  document: object | null; // Retrieved fields, or null if not stored
  highlights: Record<string, string[]>; // Highlighted fragments per requested field
}

highlights maps each field named in highlight.fields to its highlighted fragments (best first); a field that did not highlight is absent from the object, and highlights is {} when highlight was not requested. See Highlighting.


Fusion algorithms

RRF

new RRF(k?: number)  // default 60.0

Reciprocal Rank Fusion. Merges lexical and vector result lists by rank position.

WeightedSum

new WeightedSum(
  lexicalWeight?: number,  // default 0.5
  vectorWeight?: number,   // default 0.5
)

Normalises both score lists independently, then combines them.


Text analysis

SynonymDictionary

class SynonymDictionary {
  constructor();
  addSynonymGroup(terms: string[]): void;
}

WhitespaceTokenizer

class WhitespaceTokenizer {
  constructor();
  tokenize(text: string): Token[];
}

SynonymGraphFilter

class SynonymGraphFilter {
  constructor(
    dictionary: SynonymDictionary,
    keepOriginal?: boolean,  // default true
    boost?: number,          // default 1.0
  );
  apply(tokens: Token[]): Token[];
}

Token

interface Token {
  text: string;
  position: number;
  startOffset: number;
  endOffset: number;
  boost: number;
  stopped: boolean;
  positionIncrement: number;
  positionLength: number;
  tokenType?: string;
}

startOffset and endOffset are UTF-8 byte offsets into the original text. For non-ASCII text they are not JavaScript string (UTF-16) indices; slice the encoded text instead: new TextDecoder().decode(new TextEncoder().encode(text).slice(tok.startOffset, tok.endOffset)).

tokenType is one of "alphanum", "num", "cjk", "katakana", "hiragana", "hangul", "punctuation", "whitespace", "synonym", "email", "url" and "other". SynonymGraphFilter.apply keeps each token’s offsets and type, and gives each synonym it inserts the type "synonym" and the offsets of the words it replaces. It throws on any other tokenType.

A token built by hand may leave tokenType out. A multi-word synonym then matches only if the words’ offsets touch, so set tokenType: "alphanum" on words separated by spaces.


Field value types

JavaScript values are automatically converted to Laurus DataValue types:

JavaScript typeLaurus typeNotes
nullNull
booleanBool
number (integer)Int64
number (float)Float64
stringTextISO 8601 strings become DateTime
number[] (all integers)Int64ArrayMulti-valued integer field; vector fields cast the array to f32. An empty array is an empty Int64Array
number[]Float64ArrayMulti-valued float field (integers widened); vector fields cast the array to f32
{ lat, lon }GeoTwo number values
{ x, y, z }GeoEcefThree number values, meters (3D ECEF Cartesian)
{ lat, lon }[]GeoArrayArray of { lat, lon } objects; requires multiValued: true on the field
{ x, y, z }[]GeoEcefArrayArray of { x, y, z } objects; requires multiValued: true on the field
number[][]VectorArrayToken vectors of a multi-vector field (addMultiVectorField): 1 to 8,192 arrays of the field’s dimension. All inner arrays must have the same length, otherwise ingestion rejects with token vectors must share one dimension. A field with a token-level embedder also accepts a string, which is embedded into token vectors. Not stored, so absent from getDocuments and search results
string[] (all RFC 3339)DateTimeArrayArray of RFC 3339 datetime strings (same infer_from_json rules as the HTTP gateway); requires multiValued: true on the field
string[] (not all RFC 3339)TextArrayArray of strings (Issue #1175); requires multiValued: true on the field. Read back as string[]. On a declared multi-valued Bytes field, the same array of base64 strings is instead decoded element-wise into BytesArray by the schema-aware coercion (Issue #1176)
boolean[]BoolArrayArray of booleans (same infer_from_json rules as the HTTP gateway); requires multiValued: true on the field. A mixed array such as [true, 1] is rejected