API Reference
Index
The primary entry point. Wraps the Laurus search engine.
Laurus::Index.new(path: nil, schema: nil, wal_sync_policy: nil, commit_policy: nil)
Constructor
| Parameter | Type | Default | Description |
|---|---|---|---|
path: | String | nil | nil | Directory path for persistent storage. nil creates an in-memory index. When given, the directory follows the <path>/schema.toml + <path>/store/ layout laurus-cli create index/--index-dir uses, so an index built here can also be opened by the CLI (and vice versa). See below. |
schema: | Schema | nil | nil | Schema definition. Only meaningful when creating a new file-backed index (or an in-memory one); must be omitted when reopening an existing file-backed index – the persisted schema is loaded instead. An empty schema is used when omitted for a new index. |
wal_sync_policy: | WalSyncPolicy | nil | nil | Write-ahead log (WAL) durability policy. nil keeps the default per-record fsync. See WAL sync policy & durability. |
commit_policy: | CommitPolicy | nil | nil | Auto-commit policy. nil keeps the default manual mode (the caller drives every commit). See Commit policy & auto-commit. |
Creating vs. reopening a file-backed index (path: given): if <path>/schema.toml does not yet exist, this call creates a new index and persists schema: (or an empty schema, if omitted) to it. If <path>/schema.toml already exists, this call reopens the index – schema: must be omitted, or an ArgumentError is raised (it would be ambiguous which schema should win). An ArgumentError is also raised if path: contains an index in the layout used before this convention was introduced (segment files directly under path:, no schema.toml).
Methods
| Method | Description |
|---|---|
put_document(id, doc) | Upsert a document. Replaces all existing versions with the same ID. |
add_document(id, doc) | Append a document chunk without removing existing versions. |
put_documents(docs) | Batched upsert. docs is an Array of [id, hash] pairs, applied in order with one WAL fsync per batch (duplicate ids dedup, last wins). Fails fast at the first bad entry; the applied prefix is not rolled back. |
add_documents(docs) | Batched chunk append. Like put_documents but repeated ids accumulate as separate versions. |
get_documents(id) -> Array<Hash> | Return all stored versions for the given ID. |
delete_documents(id) | Delete all versions for the given ID. |
commit | Flush buffered writes and make all pending changes searchable. |
flush_wal | Force a durable WAL barrier on demand. Synchronously fsyncs any unsynced WAL records and returns nil. Useful when running under a group-commit policy (see below). |
search(query, limit: 10, offset: 0, highlight: nil, rescore: nil) -> Array<SearchResult> | Execute a search query. rescore: takes a LateInteractionRescore that reorders the top results (Issue #1351); any other value raises TypeError. See LateInteractionRescore. |
search_batch(queries, limit: 10, offset: 0, highlight: nil) -> Array<Array<SearchResult>> | Execute multiple independent searches in one call. Each query is dispatched in parallel on the underlying tokio runtime. results[i] corresponds to queries[i]. Empty input returns []. highlight: applies identically to every query in the batch. There is no rescore: keyword; a SearchRequest element is rescored by its own rescore:. |
stats -> Hash | Return index statistics ("document_count", "vector_fields"). |
search query argument
The query parameter accepts any of the following:
- A DSL string (e.g.
"title:hello","embedding:\"memory safety\"") - A lexical query object (
TermQuery,PhraseQuery,BooleanQuery, …) - A vector query object (
VectorQuery,VectorTextQuery) - A
SearchRequestfor full control
When query is a SearchRequest, its own limit:, offset:, highlight: and rescore: are used and the keyword arguments of search are ignored.
The same value kinds are accepted as the elements of search_batch’s queries Array — DSL strings, query objects, and SearchRequest instances may be mixed within a single batch.
Highlighting
search/search_batch’s highlight: keyword (Issue #1134) requests highlighted fragments per field on each hit’s SearchResult#highlights. It accepts:
- An Array of field names:
highlight: ["body"] - A Hash with a required
fieldskey (String or Symbol) plus any of the optionalHighlightConfigknobs —max_fragments,fragment_size,tag,css_class,require_field_match,max_analyzed_chars,return_entire_field_if_no_highlight(String or Symbol keys both work) — e.g.highlight: {fields: ["body"], tag: "em", max_fragments: 2}
results = index.search("body:rust", highlight: ["body"])
results[0].highlights # => {"body" => ["<mark>Rust</mark> is a systems programming language."]}
Highlighting follows the query passed to search/search_batch (or SearchRequest’s query/lexical_query — see below), and only stored: true text fields can be highlighted; a field that is absent, not stored, or not text is silently skipped. Omitting highlight: leaves every result’s highlights empty. The same highlight: keyword is also accepted by SearchRequest.new.
WAL sync policy & durability
The write-ahead log (WAL) protects committed data against crashes. By
default the WAL is fully durable: every record is fsynced before the write
returns. You can trade some durability for higher write throughput by opting
into group commit, which batches fsync calls.
WalSyncPolicy
Laurus::WalSyncPolicy is an immutable value object describing how the WAL
is flushed. Pass it to Index.new(wal_sync_policy:).
# Default: durable per write (each record is fsynced individually).
Laurus::WalSyncPolicy.per_record
# Group commit: batch fsyncs to amortise their cost.
Laurus::WalSyncPolicy.group(
max_records: nil, # flush after this many records (default 1024)
max_bytes: nil, # flush after this many bytes (default 1 MiB)
max_interval_ms: nil, # also flush periodically after this many ms
)
| Constructor | Description |
|---|---|
WalSyncPolicy.per_record | Default. Every record is fsynced before the write returns — fully durable per write. |
WalSyncPolicy.group(max_records:, max_bytes:, max_interval_ms:) | Batch fsyncs. With no arguments uses the defaults (max_records: 1024, max_bytes: 1 MiB, no timer). The WAL is flushed when either max_records or max_bytes accumulate, and on every commit. Pass max_interval_ms: to also flush on a periodic timer. |
Group commit is analogous to SQLite’s synchronous = NORMAL: a crash can lose
at most the last unsynced batch of records, but the index never corrupts.
Records are always made durable at commit, so a successful commit is a
durability barrier regardless of policy.
Forcing a flush
Call flush_wal to force a durable barrier between commits — for example
before signalling that a batch has been safely persisted. It synchronously
fsyncs any unsynced records and returns nil. Under the default
per-record policy it is effectively a no-op.
# Opt into group commit, then force durability on demand.
policy = Laurus::WalSyncPolicy.group(max_records: 4096, max_bytes: 4 * 1024 * 1024)
index = Laurus::Index.new(path: "./myindex", wal_sync_policy: policy)
index.put_document("doc1", { "title" => "Hello" })
index.flush_wal # records persisted even though the group batch is not full
Commit policy & auto-commit
By default the caller drives every commit: buffered writes only become
searchable once you call commit explicitly. You can hand that responsibility
to the engine with an auto-commit policy, which commits automatically after
a fixed number of applied documents or on a periodic timer.
CommitPolicy
Laurus::CommitPolicy is an immutable value object describing when the engine
materialises buffered writes into the stores. Pass it to
Index.new(commit_policy:).
# Default: manual — the caller drives every commit.
Laurus::CommitPolicy.manual
# Auto-commit: commit after every N applied documents.
Laurus::CommitPolicy.every_docs(1000)
# Auto-commit: commit at least every N milliseconds (native only).
Laurus::CommitPolicy.interval_ms(5000)
| Constructor | Description |
|---|---|
CommitPolicy.manual | Default. No auto-commit — the caller drives every commit. |
CommitPolicy.every_docs(n) | Auto-commit after every n applied documents. Counted across both singular and batch ingest, including every n documents within a single batch. |
CommitPolicy.interval_ms(ms) | Auto-commit at least every ms milliseconds via a background timer, so a trailing partial batch is committed even while ingestion is idle. The time-based counterpart of every_docs. Default: none. Native only — under WebAssembly (wasm32) there are no background threads, so the engine treats it as a no-op (the value still constructs, but no timed commit happens). |
every_docs(0) is valid and disables auto-commit, making it equivalent to
manual.
Commit policy is orthogonal to WalSyncPolicy:
WalSyncPolicy governs WAL fsync durability, whereas CommitPolicy governs
when the stores materialise buffered writes. Set them independently.
# Auto-commit after every 1000 applied documents.
policy = Laurus::CommitPolicy.every_docs(1000)
index = Laurus::Index.new(path: "./myindex", commit_policy: policy)
Schema
Defines the fields and index types for an Index.
Laurus::Schema.new
Field methods
| Method | Description |
|---|---|
add_text_field(name, stored: true, indexed: true, term_vectors: true, doc_values: true, analyzer: nil, multi_valued: false, position_increment_gap: 100) | Full-text field (inverted index, BM25). term_vectors: controls whether term positions are stored, read by phrase and span queries. doc_values: controls whether the value is also copied into DocValues, the column-oriented store sort/facet/aggregation read from (Issue #1047); takes effect only when stored: true. Pass multi_valued: true to accept an Array of Strings (Issue #1175): a term query matches if any element contains the term, a phrase query never spans two elements unless its slop reaches position_increment_gap: (default 100; 0 numbers the elements as if concatenated), and values are read back as an Array of Strings. analyzer: is the name of a parameter-less built-in ("standard", "english", "keyword", "simple", "noop") or a custom name registered via add_analyzer. The Japanese preset requires a Lindera dictionary path, so register it as a custom analyzer with a lindera tokenizer and reference it by name. |
add_integer_field(name, stored: true, indexed: true, multi_valued: false, doc_values: true) | 64-bit integer field. Pass multi_valued: true to accept arrays of integers (range queries match if any value satisfies the predicate). See doc_values: above. |
add_float_field(name, stored: true, indexed: true, multi_valued: false, doc_values: true) | 64-bit float field. Pass multi_valued: true to accept arrays of floats (range queries match if any value satisfies the predicate). See doc_values: above. |
add_boolean_field(name, stored: true, indexed: true, multi_valued: false, doc_values: true) | Boolean field. Pass multi_valued: true to accept an Array of true / false values (a term query such as flags:true matches if any element equals the value; values are read back as an Array of true / false). See doc_values: above. |
add_bytes_field(name, stored: true, multi_valued: false) | Raw bytes field. No doc_values: option: a Bytes value is never written to DocValues regardless. Pass multi_valued: true to accept an Array of base64 Strings (Issue #1176), decoded element-wise the same way a single base64 String is; Bytes is never indexed, so unlike every other multi_valued: option this has no query-matching semantics — it only governs the stored shape and ingestion arity. Values are read back as an Array of (binary) Strings, mirroring the scalar field. |
add_geo_field(name, stored: true, indexed: true, multi_valued: false, doc_values: true) | Geographic coordinate field (lat/lon). Pass multi_valued: true to accept an Array of { "lat" => .., "lon" => .. } Hashes (distance / bounding-box queries match if any point satisfies the predicate). See doc_values: above. |
add_geo3d_field(name, stored: true, indexed: true, multi_valued: false, doc_values: true) | 3D ECEF Cartesian point field (x, y, z in metres). Pass multi_valued: true to accept an Array of { "x" => .., "y" => .., "z" => .. } Hashes (distance / bounding-box / nearest queries match if any point satisfies the predicate). See Geo3d concepts and doc_values: above. |
add_datetime_field(name, stored: true, indexed: true, multi_valued: false, doc_values: true) | UTC datetime field. Pass multi_valued: true to accept an Array of Time / DateTime / RFC 3339 String values (range queries match if any instant satisfies the predicate; values are read back as an Array of RFC 3339 Strings in UTC). See doc_values: above. |
add_hnsw_field(name, dimension, distance: "cosine", m: 16, ef_construction: 200, quantizer: nil, subvector_count: nil, rerank_storage: nil, embedder: nil, pq_codebook_path: nil, base_weight: 1.0) | HNSW approximate nearest-neighbor vector field. base_weight sets this field’s relative scoring priority when searched alongside other vector fields (Issue #1084); see Vector Search → Weights. |
add_flat_field(name, dimension, distance: "cosine", embedder: nil, base_weight: 1.0) | Flat (brute-force) vector field. |
add_ivf_field(name, dimension, distance: "cosine", n_clusters: 100, n_probe: 1, embedder: nil, base_weight: 1.0) | IVF approximate nearest-neighbor vector field. |
add_multi_vector_field(name, dimension, distance: "cosine", storage: "f32", embedder: nil) | Multi-vector field holding a variable number of token vectors per document, for example ColBERT per-token embeddings (Issue #1351). Its value is an Array of numeric Arrays, one per token, each of dimension elements. It has no ANN index and is read only by the late-interaction rescore (see LateInteractionRescore); its token vectors are not stored, so get_documents and search results never include the field. distance: must be "cosine" or "dot_product" and dimension greater than 0, both checked when the field is added (ArgumentError). storage: sets the on-disk element kind of each token vector — "f32" (default, exact), "f16" (2x smaller), or "int8" (~4x smaller, lossy quantization) — also checked when the field is added (ArgumentError). embedder: names a token-level embedder (a "candle_colbert" one, see Embedder types), which embeds text values and rescore query text. See Multi-Vector Fields. |
Vector quantization & rerank storage (HNSW fields):
quantizer—"scalar_8bit"(default, 4× compression) or"product_quantization"for higher compression. Product quantization requiressubvector_count(must dividedimension).rerank_storage— set to"f32"to write a full-precision*.hnsw.f32sidecar enabling exact Stage-2 rerank; omit to keep the int8-only segment.pq_codebook_path— storage-relative file name of a shared PQ codebook (Issue #631), trained once via thelaurus train pq-codebookCLI command. Only meaningful withquantizer: "product_quantization"; commits then encode against the pre-trained codebook instead of re-training k-means per segment. Omit to keep per-segment training.
Every add_*_field method above raises ArgumentError when name starts with _ (other than _id), and adds nothing. A schema loaded with from_toml / from_toml_file keeps accepting such a field, so a persisted schema still loads; creating a new Index from it raises ArgumentError. See Field Naming.
Other methods
| Method | Description |
|---|---|
add_embedder(name, config) | Register a named embedder definition. config is a Hash with a "type" key, with either String or Symbol keys (see below). A missing or unknown type, or a missing required key, raises ArgumentError (invalid embedder config: ...). |
add_analyzer(name, tokenizer, char_filters: nil, token_filters: nil) | Register a custom analyzer definition. tokenizer is required; char_filters:/token_filters: are optional Arrays of Hashes. Each Hash uses the same {type: "..."} shape as the schema TOML/JSON format (see below), with either String or Symbol keys. A name reserved for a built-in analyzer (standard, keyword, english, simple, noop) raises ArgumentError, and so does creating a new Index from a schema that defines one (e.g. loaded with from_toml). Semantic validity (e.g. a malformed regex) is checked when the schema is used to build an Index, not here. |
analyzer_names -> Array<String> | Return the names of custom analyzers registered via add_analyzer or loaded from TOML. |
Laurus::Schema.from_toml(toml_str) -> Schema (class method) | Parse a schema from a TOML string, in the same format laurus-cli create index --schema accepts. |
Laurus::Schema.from_toml_file(path) -> Schema (class method) | Load a schema from a TOML file. |
to_toml -> String | Serialize this schema to a TOML string in the same format laurus-cli accepts. |
to_toml_file(path) | Write this schema to a TOML file. |
set_default_fields(fields) | Set the default fields used when no field is specified in a query. fields is an Array of Strings. |
set_dynamic_field_policy(policy) | Set how undeclared fields are handled. policy is "strict", "dynamic" (default), or "ignore". See notes below. |
dynamic_field_policy -> String | Return the current policy as a lowercase string. |
field_names -> Array<String> | Return the list of field names defined in this schema. |
Dynamic field policy
Controls what happens when a document is ingested with field names that are not declared in the schema:
"strict"— Reject the document."dynamic"(default) — Infer a type for each undeclared field and add it to the schema. Warning: integer fields silently truncate incoming float values (3.14→3). Use"strict"if you need to reject such type mismatches."ignore"— Silently drop the undeclared fields.
See Schema & Fields for the full behaviour matrix.
Embedder types
See Schema Format Reference → Embedders for the canonical description of each type.
"type" | Required keys | Feature flag |
|---|---|---|
"precomputed" | – | (always available) |
"candle_bert" | "model" | embeddings-candle |
"candle_clip" | "model" | embeddings-multimodal |
"openai" | "model" | embeddings-openai |
"candle_colbert" | "model" | embeddings-candle |
"candle_colbert" also takes the optional keys "revision", "query_maxlen" and "doc_maxlen". It produces token vectors, so only a multi-vector field (add_multi_vector_field) can use it. Keys may be Strings or Symbols:
schema = Laurus::Schema.new
schema.add_embedder(
"colbert",
{ type: "candle_colbert", model: "colbert-ir/colbertv2.0", query_maxlen: 32 },
)
schema.add_multi_vector_field("body_colbert", 128, embedder: "colbert")
Analyzer components
Used by add_analyzer(name, tokenizer, char_filters: nil, token_filters: nil)
and by the [analyzers.<name>] TOML section. tokenizer is a single Hash;
char_filters:/token_filters: are Arrays of Hashes, applied in Array order.
See Schema Format Reference → Analyzers for the canonical description of each component.
Tokenizers (tokenizer, exactly one):
type | Required keys | Optional keys |
|---|---|---|
"whitespace" | – | – |
"unicode_word" | – | – |
"regex" | – | pattern (default \w+), gaps (default false) |
"ngram" | min_gram, max_gram | – |
"lindera" | mode, dict | user_dict |
"whole" | – | – |
Char filters (char_filters, applied to raw text before tokenization):
type | Required keys | Optional keys |
|---|---|---|
"unicode_normalization" | form ("nfc"/"nfd"/"nfkc"/"nfkd") | – |
"pattern_replace" | pattern, replacement | – |
"mapping" | mapping (Hash of string replacements) | – |
"japanese_iteration_mark" | – | kanji (default true), kana (default true) |
Token filters (token_filters, applied to the token stream after tokenization):
type | Required keys | Optional keys |
|---|---|---|
"lowercase" | – | – |
"stop" | – | words (default: English stop words) |
"stem" | – | stem_type ("porter"/"simple"/"identity") |
"boost" | boost | – |
"limit" | limit | – |
"strip" | – | – |
"remove_empty" | – | – |
"flatten_graph" | – | – |
schema = Laurus::Schema.new
schema.add_analyzer(
"ja_ipadic",
{ type: "lindera", mode: "normal", dict: "/var/lib/lindera/ipadic" },
char_filters: [
{ type: "unicode_normalization", form: "nfkc" },
{ type: "japanese_iteration_mark" },
],
token_filters: [{ type: "lowercase" }],
)
schema.add_text_field("title", analyzer: "ja_ipadic")
Distance metrics
| Value | Description |
|---|---|
"cosine" | Cosine similarity (default) |
"euclidean" | Euclidean distance |
"dot_product" | Dot product |
"manhattan" | Manhattan distance |
"angular" | Angular distance |
Query classes
TermQuery
Laurus::TermQuery.new(field, term)
Matches documents containing the exact term in the given field.
PhraseQuery
Laurus::PhraseQuery.new(field, terms)
Matches documents containing the terms in order. terms is an Array of Strings.
FuzzyQuery
Laurus::FuzzyQuery.new(field, term, max_edits: 2)
Approximate match allowing up to max_edits edit-distance errors.
WildcardQuery
Laurus::WildcardQuery.new(field, pattern)
Pattern match. * matches any sequence of characters, ? matches any single character.
NumericRangeQuery
Laurus::NumericRangeQuery.new(field, min: nil, max: nil)
Matches numeric values in the range [min, max]. Pass nil for an open bound. The type (integer or float) is inferred from the Ruby type of min/max.
DateTimeRangeQuery
Laurus::DateTimeRangeQuery.new(field, min: nil, max: nil)
Matches DateTime values in the range [min, max] (both bounds inclusive). Pass nil (or omit the keyword) for an open bound. A bound is a String literal in any form the query DSL accepts — RFC 3339 ("2024-01-01T09:00:00+09:00", normalized to UTC), naive "YYYY-MM-DDTHH:MM:SS[.fff]" (UTC), or "YYYY-MM-DD" (midnight UTC) — or any object responding to iso8601 (Time, DateTime). A malformed bound raises ArgumentError at construction. Usable wherever a query object is accepted (Index#search, BooleanQuery, SearchRequest).
GeoDistanceQuery
Laurus::GeoDistanceQuery.within_radius(field, lat, lon, distance_m)
Geo-distance (radius) search. Returns documents whose (lat, lon) coordinate
is within distance_m metres of the given point.
GeoBoundingBoxQuery
Laurus::GeoBoundingBoxQuery.within_bounding_box(
field, min_lat, min_lon, max_lat, max_lon,
)
Geo bounding-box search. Returns documents whose (lat, lon) coordinate lies
inside the axis-aligned [min_lat, max_lat] × [min_lon, max_lon] rectangle.
Geo3dDistanceQuery
Laurus::Geo3dDistanceQuery.within_sphere(field, x, y, z, distance_m)
Sphere search over a 3D ECEF point field. Returns documents whose (x, y, z)
coordinate is within distance_m metres of the centre. See
Geo3d concepts for ECEF theory.
Geo3dBoundingBoxQuery
Laurus::Geo3dBoundingBoxQuery.within_box(
field,
min_x, min_y, min_z,
max_x, max_y, max_z,
)
Axis-aligned 3D bounding-box search.
Geo3dNearestQuery
Laurus::Geo3dNearestQuery.k_nearest(
field, x, y, z, k,
initial_radius_m: nil,
max_radius_m: nil,
)
k-nearest-neighbour search over a 3D ECEF point field. The optional
initial_radius_m: and max_radius_m: keyword arguments tune the
iterative-expansion search cone.
BooleanQuery
bq = Laurus::BooleanQuery.new
bq.must(query)
bq.should(query)
bq.must_not(query)
Compound boolean query. must clauses all have to match; must_not clauses must not match. should clauses contribute to scoring; at least one of them must match if there are no must clauses.
SpanQuery
# Single term
Laurus::SpanQuery.term(field, term)
# Near: terms within slop positions
Laurus::SpanQuery.near(field, terms, slop: 0, ordered: true)
# Near with nested SpanQuery clauses
Laurus::SpanQuery.near_spans(field, clauses, slop: 0, ordered: true)
# Containing: big span contains little span
Laurus::SpanQuery.containing(field, big, little)
# Within: include span within exclude span at max distance
Laurus::SpanQuery.within(field, include_span, exclude_span, distance)
Positional / proximity span queries. near takes an Array of term Strings, while near_spans takes an Array of SpanQuery objects for nested expressions.
VectorQuery
Laurus::VectorQuery.new(field, vector)
Approximate nearest-neighbor search using a pre-computed embedding vector. vector is an Array of Floats.
VectorTextQuery
Laurus::VectorTextQuery.new(field, text)
Converts text to an embedding at query time and runs vector search. Requires an embedder configured on the index.
SearchRequest
Full-featured search request for advanced control.
Laurus::SearchRequest.new(
query: nil,
lexical_query: nil,
vector_query: nil,
filter_query: nil,
fusion: nil,
limit: 10,
offset: 0,
highlight: nil,
rescore: nil,
)
| Parameter | Description |
|---|---|
query: | A DSL string or single query object. Mutually exclusive with lexical_query: / vector_query:. |
lexical_query: | Lexical component for explicit hybrid search. |
vector_query: | Vector component for explicit hybrid search. |
filter_query: | Lexical filter applied after scoring. |
fusion: | Fusion algorithm (RRF or WeightedSum). Defaults to RRF(k: 60) when both components are set. |
limit: | Maximum number of results (default 10). |
offset: | Pagination offset (default 0). |
highlight: | Same Array-or-Hash shape as Index#search’s highlight: (Issue #1134). See Highlighting. |
rescore: | A LateInteractionRescore that reorders the top results of the request — lexical, vector, or hybrid (Issue #1351); any other value raises TypeError. See LateInteractionRescore. |
LateInteractionRescore
Late-interaction (ColBERT MaxSim) rescore of the top search results (Issue #1351). Pass it as rescore: to Index#search or SearchRequest.new.
Laurus::LateInteractionRescore.new(field, query, window_size: nil)
| Parameter | Type | Default | Description |
|---|---|---|---|
field | String | – | A multi-vector field (add_multi_vector_field). |
query | String | Array<Array<Numeric>> | – | Query text, embedded by the field’s token-level embedder (a "candle_colbert" one), or the query’s token vectors, computed with the same model as the documents’ token vectors. Integer elements are accepted. |
window_size: | Integer | nil | nil (100) | How many top first-stage results to rescore, at most 10,000. nil keeps the default of 100. |
| Method | Description |
|---|---|
window_size -> Integer | Return how many top results are rescored. |
inspect -> String | Return a summary such as LateInteractionRescore(field="tokens", window_size=100). |
The top window_size first-stage results are reordered by their MaxSim against the field, and a rescored result’s score is its MaxSim. Results beyond the window keep their first-stage order and score, after the rescored ones. See Vector Search → Late-Interaction Rescore for the ordering rules and the similarity.
Errors: the constructor raises TypeError when query is neither a String nor an Array of numeric Arrays. The other values are checked by the engine when searching, which raises ArgumentError with a message containing rescore: ... when the field is unknown or not a multi-vector field, a text query is blank or the field has no token-level embedder, the query does not hold between 1 and 1,024 token vectors of the field’s dimension, or window_size is not between 1 and 10,000.
schema = Laurus::Schema.new
schema.add_text_field("title")
schema.add_multi_vector_field("tokens", 2, distance: "dot_product")
index = Laurus::Index.new(schema: schema)
index.put_document("a", { "title" => "rust", "tokens" => [[0.1, 0.0]] })
index.put_document("b", { "title" => "rust language", "tokens" => [[0.9, 0.2]] })
index.commit
rescore = Laurus::LateInteractionRescore.new("tokens", [[1.0, 0.0], [0.0, 1.0]], window_size: 50)
results = index.search("title:rust", rescore: rescore)
results.map(&:id) # => ["b", "a"] -- "b" has MaxSim 1.1, "a" 0.1
# With a "candle_colbert" embedder on the field, the query can be text.
rescore = Laurus::LateInteractionRescore.new("body_colbert", "how do lifetimes work")
SearchResult
Returned by Index#search.
result.id # => String -- External document identifier
result.score # => Float -- Relevance score
result.document # => Hash|nil -- Retrieved field values, or nil if deleted
result.highlights # => Hash -- Highlighted fragments per requested field
highlights maps each field named in highlight: to its highlighted fragments (best first); a field that did not highlight is absent from the Hash, and highlights is {} when highlight: was not requested. See Highlighting.
Fusion algorithms
RRF
Laurus::RRF.new(k: 60.0)
Reciprocal Rank Fusion. Merges lexical and vector result lists by rank position. k is a smoothing constant; higher values reduce the influence of top-ranked results.
WeightedSum
Laurus::WeightedSum.new(lexical_weight: 0.5, vector_weight: 0.5)
Normalises both score lists independently, then combines them as lexical_weight * lexical_score + vector_weight * vector_score.
Text analysis
SynonymDictionary
dict = Laurus::SynonymDictionary.new
dict.add_synonym_group(["fast", "quick", "rapid"])
A dictionary of synonym groups. All terms in a group are treated as synonyms of each other.
WhitespaceTokenizer
tokenizer = Laurus::WhitespaceTokenizer.new
tokens = tokenizer.tokenize("hello world")
Splits text on whitespace boundaries and returns an Array of Token objects.
SynonymGraphFilter
filter = Laurus::SynonymGraphFilter.new(dictionary, keep_original: true, boost: 1.0)
expanded = filter.apply(tokens)
Token filter that expands tokens with their synonyms from a SynonymDictionary.
Token
token.text # => String -- The token text
token.position # => Integer -- Position in the token stream
token.start_offset # => Integer -- UTF-8 byte start offset in the original text
token.end_offset # => Integer -- UTF-8 byte end offset in the original text
token.boost # => Float -- Score boost factor (1.0 = no adjustment)
token.stopped # => Boolean -- Whether removed by a stop filter
token.position_increment # => Integer -- Difference from the previous token's position
token.position_length # => Integer -- Number of positions spanned
token.token_type # => String or nil -- Token type, e.g. "alphanum"
The offsets count bytes, not characters; for non-ASCII text slice the bytes: text.byteslice(token.start_offset...token.end_offset).
token_type is one of "alphanum", "num", "cjk", "katakana", "hiragana", "hangul", "punctuation", "whitespace", "synonym", "email", "url" and "other", or nil. SynonymGraphFilter#apply keeps each token’s offsets and type, and gives each synonym it inserts the type "synonym" and the offsets of the words it replaces.
Field value types
Ruby values are automatically converted to Laurus DataValue types:
| Ruby type | Laurus type | Notes |
|---|---|---|
nil | Null | |
true / false | Bool | |
Integer | Int64 | |
Float | Float64 | |
String | Text | |
Array of Integer | Int64Array | Multi-valued integer field; vector fields cast the array to f32. An empty Array is an empty Int64Array |
Array of numerics | Float64Array | Multi-valued float field (integers widened); vector fields cast the array to f32 |
Array of numeric Arrays | VectorArray | Token vectors of a multi-vector field (add_multi_vector_field), one Array per token, each of the field’s dimension (Issue #1351). Integer elements are accepted; any other element, such as true or a String, raises TypeError. Not returned by get_documents or search results |
Hash with "lat", "lon" | Geo | Two Float values |
Hash with "x", "y", "z" | GeoEcef | Three Float values, meters (3D ECEF Cartesian) |
Array of Hash with "lat", "lon" | GeoArray | Requires multi_valued: true on the field |
Array of Hash with "x", "y", "z" | GeoEcefArray | Requires multi_valued: true on the field |
Time / String responding to iso8601 | DateTime | Converted via iso8601 |
Array containing Time (or other objects responding to iso8601) | DateTimeArray | Each element parsed as RFC 3339 (objects responding to iso8601 are converted first); a non-datetime element raises ArgumentError. Requires multi_valued: true on the field |
Array of String | DateTimeArray or TextArray | A DateTimeArray when every element parses as RFC 3339, a TextArray otherwise (Issue #1175; read back as an Array of Strings). Requires multi_valued: true on the field. On a declared multi-valued Bytes field, the same Array of base64 Strings is instead decoded element-wise (Issue #1176) |
Array of true / false | BoolArray | All elements must be true or false; a mixed Array such as [true, 1] raises TypeError. Requires multi_valued: true on the field |