API Reference
Index
The primary entry point. Wraps the Laurus search engine.
class Index:
def __init__(
self,
path: str | None = None,
schema: Schema | None = None,
wal_sync_policy: WalSyncPolicy | None = None,
commit_policy: CommitPolicy | None = None,
) -> None: ...
Constructor
| Parameter | Type | Default | Description |
|---|---|---|---|
path | str | None | None | Directory path for persistent storage. None creates an in-memory index. When given, the directory follows the <path>/schema.toml + <path>/store/ layout laurus-cli create index/--index-dir uses, so an index built here can also be opened by the CLI (and vice versa). See below. |
schema | Schema | None | None | Schema definition. Only meaningful when creating a new file-backed index (or an in-memory one); must be omitted (None) when reopening an existing file-backed index — the persisted schema is loaded instead. An empty schema is used when omitted for a new index. |
wal_sync_policy | WalSyncPolicy | None | None | WAL durability policy. None keeps the default per-record fsync. See WAL sync policy / durability. |
commit_policy | CommitPolicy | None | None | Auto-commit policy. None keeps the default manual commit (no auto-commit). See Commit policy / auto-commit. |
Creating vs. reopening a file-backed index (path given): if <path>/schema.toml does not yet exist, this call creates a new index and persists schema (or an empty schema, if omitted) to it. If <path>/schema.toml already exists, this call reopens the index — schema must be omitted, or a ValueError is raised (it would be ambiguous which schema should win). A ValueError is also raised if path contains an index in the layout used before this convention was introduced (segment files directly under path, no schema.toml).
Methods
| Method | Description |
|---|---|
put_document(id, doc) | Upsert a document. Replaces all existing versions with the same ID. |
add_document(id, doc) | Append a document chunk without removing existing versions. |
put_documents(docs) | Batched upsert. docs is an iterable of (id, dict) pairs, applied in order with one WAL fsync per batch (duplicate ids dedup, last wins). Fails fast at the first bad entry; the applied prefix is not rolled back. |
add_documents(docs) | Batched chunk append. Like put_documents but repeated ids accumulate as separate versions. |
get_documents(id) -> list[dict] | Return all stored versions for the given ID. |
delete_documents(id) | Delete all versions for the given ID. |
commit() | Flush buffered writes and make all pending changes searchable. |
flush_wal() | Force a durable WAL barrier. See WAL sync policy / durability. |
search(query, *, limit=10, offset=0, highlight=None, rescore=None) -> list[SearchResult] | Execute a search query. rescore takes a LateInteractionRescore that reorders the top results by late interaction over a multi-vector field (Issue #1351). When query is a SearchRequest, the request’s own rescore is used and this keyword is ignored. |
search_batch(queries, *, limit=10, offset=0, highlight=None) -> list[list[SearchResult]] | Execute multiple independent searches in one call. Each query is dispatched in parallel on the underlying tokio runtime. results[i] corresponds to queries[i]. Empty input returns []. highlight applies identically to every query in the batch. There is no rescore keyword; pass a SearchRequest with rescore= for each query that needs one. |
stats() -> dict | Return index statistics (document_count, vector_fields). |
search query argument
The query parameter accepts any of the following:
- A DSL string (e.g.
"title:hello","embedding:\"memory safety\"") - A lexical query object (
TermQuery,PhraseQuery,BooleanQuery, …) - A vector query object (
VectorQuery,VectorTextQuery) - A
SearchRequestfor full control
The same value kinds are accepted as the elements of search_batch’s queries list — DSL strings, query objects, and SearchRequest instances may be mixed within a single batch.
Highlighting
search/search_batch’s highlight parameter (Issue #1134) requests highlighted fragments per field on each hit’s SearchResult.highlights. It accepts:
- A list of field names:
highlight=["body"] - A dict with a required
"fields"key plus any of the optionalHighlightConfigknobs:max_fragments,fragment_size,tag,css_class,require_field_match,max_analyzed_chars,return_entire_field_if_no_highlight— e.g.highlight={"fields": ["body"], "tag": "em", "max_fragments": 2}
results = index.search("body:rust", highlight=["body"])
print(results[0].highlights) # {"body": ["<mark>Rust</mark> is a systems programming language."]}
Highlighting follows the query passed to search/search_batch (or SearchRequest.query/lexical_query — see below), and only stored: true text fields can be highlighted; a field that is absent, not stored, or not text is silently skipped. Omitting highlight leaves every result’s highlights empty. The same highlight= keyword is also accepted by SearchRequest.
WAL sync policy / durability
For a persistent index, every write is appended to a write-ahead log (WAL).
By default the WAL is fsync-ed on every record, so each write is fully
durable as soon as the call returns. The constructor accepts an optional
wal_sync_policy to trade some durability for higher write throughput, and
flush_wal() forces a durable barrier on demand.
class WalSyncPolicy:
@staticmethod
def per_record() -> WalSyncPolicy: ...
@staticmethod
def group(
max_records: int | None = None,
max_bytes: int | None = None,
max_interval_ms: int | None = None,
) -> WalSyncPolicy: ...
| Constructor | Description |
|---|---|
WalSyncPolicy.per_record() | Default. fsync after every WAL record; fully durable per write. |
WalSyncPolicy.group(...) | Group commit. Batches fsync across writes. |
group() parameters (all keyword-friendly, None keeps the default):
| Parameter | Default | Description |
|---|---|---|
max_records | 1024 | Flush once this many records have accumulated. |
max_bytes | 1048576 (1 MiB) | Flush once this many unsynced bytes have accumulated. |
max_interval_ms | None | Optional periodic flush timer (milliseconds). None disables the timer. |
With group commit the WAL is flushed when either max_records or
max_bytes is reached, and always at commit(). A crash can lose up to the
last unsynced batch — the same trade-off as SQLite’s synchronous = NORMAL.
Call flush_wal() to force everything written so far to disk without a full
commit().
| Method | Description |
|---|---|
flush_wal() | Force a durable WAL barrier now. Synchronous; returns None. |
import laurus
# Opt into group commit with a 1-second periodic flush timer.
policy = laurus.WalSyncPolicy.group(max_records=4096, max_interval_ms=1000)
index = laurus.Index(path="./myindex", wal_sync_policy=policy)
for i in range(10_000):
index.put_document(f"doc{i}", {"title": f"Document {i}"})
# Force a durable barrier without committing yet.
index.flush_wal()
index.commit() # also flushes the WAL
Omit wal_sync_policy (or pass WalSyncPolicy.per_record()) to keep the
default, fully durable behaviour.
Commit policy / auto-commit
A commit materializes buffered writes into the stores and makes all pending
changes searchable. By default an Index never auto-commits, so the caller
drives every commit() explicitly. The constructor accepts an optional
commit_policy to instead let the engine auto-commit after a fixed number of
applied documents, or at least once every fixed time interval.
class CommitPolicy:
@staticmethod
def manual() -> CommitPolicy: ...
@staticmethod
def every_docs(n: int) -> CommitPolicy: ...
@staticmethod
def interval_ms(ms: int) -> CommitPolicy: ...
| Constructor | Description |
|---|---|
CommitPolicy.manual() | Default. No auto-commit; the caller drives every commit(). |
CommitPolicy.every_docs(n) | Auto-commit after every n applied documents. |
CommitPolicy.interval_ms(ms) | Auto-commit at least every ms milliseconds via a background timer (default: none). Native-only; a no-op on wasm. |
every_docs(n) counts applied documents across both singular (put_document,
add_document) and batch (put_documents, add_documents) ingest, and
auto-commits after every n documents — including within a single batch.
every_docs(0) is valid and disables auto-commit, making it equivalent to
manual().
interval_ms(ms) is the time-based counterpart of every_docs(n): a
background timer commits at least every ms milliseconds, so a trailing
partial batch is committed even while ingestion is idle. This factory is
native-only — on wasm there are no background threads, so the engine treats
it as a no-op: the value still constructs, but no timed commit happens under
WebAssembly.
commit_policy is orthogonal to wal_sync_policy: wal_sync_policy governs
WAL fsync durability, whereas commit_policy governs when the stores
materialize pending changes into searchable state. Setting one does not affect
the other.
import laurus
# Auto-commit after every 100 applied documents.
index = laurus.Index(
path="./myindex",
commit_policy=laurus.CommitPolicy.every_docs(100),
)
for i in range(1_000):
index.put_document(f"doc{i}", {"title": f"Document {i}"})
# The engine has already committed 10 times (once per 100 documents).
Omit commit_policy (or pass CommitPolicy.manual()) to keep the default,
manual-commit behaviour.
Schema
Defines the fields and index types for an Index.
class Schema:
def __init__(self) -> None: ...
Field methods
| Method | Description |
|---|---|
add_text_field(name, *, stored=True, indexed=True, term_vectors=True, doc_values=True, multi_valued=False, position_increment_gap=100, analyzer=None) | Full-text field (inverted index, BM25). term_vectors controls whether term positions are stored, read by phrase and span queries. doc_values controls whether the value is also copied into DocValues, the column-oriented store sort/facet/aggregation read from (Issue #1047); takes effect only when stored=True. Set multi_valued=True to accept a list[str] (Issue #1175): a term query matches if any element contains the term, a phrase query never spans two elements unless its slop reaches position_increment_gap (default 100; 0 numbers the elements as if concatenated), and values are read back as a list[str]. analyzer accepts a built-in name ("standard", "english", "keyword", "simple", "noop", or any custom name registered via add_analyzer) or a dict configuring a parameterised preset such as {"language": "japanese", "mode": "normal", "dict": "/var/lib/lindera/ipadic"}. The bare string "japanese" is rejected because the preset requires a Lindera dictionary path. |
add_integer_field(name, *, stored=True, indexed=True, multi_valued=False, doc_values=True) | 64-bit integer field. Set multi_valued=True to accept arrays of integers (range queries match if any value satisfies the predicate). See doc_values above. |
add_float_field(name, *, stored=True, indexed=True, multi_valued=False, doc_values=True) | 64-bit float field. Set multi_valued=True to accept arrays of floats (range queries match if any value satisfies the predicate). See doc_values above. |
add_boolean_field(name, *, stored=True, indexed=True, multi_valued=False, doc_values=True) | Boolean field. Set multi_valued=True to accept a list[bool] (a term query such as flags:true matches if any element equals the value; values are read back as a list[bool]). See doc_values above. |
add_bytes_field(name, *, stored=True, multi_valued=False) | Raw bytes field. No doc_values option: a Bytes value is never written to DocValues regardless. Set multi_valued=True to accept a list[bytes] (Issue #1176); Bytes is never indexed, so unlike every other multi_valued option this has no query-matching semantics — it only governs the stored shape and ingestion arity. Values are read back as a list[bytes] (MIME is not preserved). |
add_geo_field(name, *, stored=True, indexed=True, multi_valued=False, doc_values=True) | Geographic coordinate field (lat/lon). Set multi_valued=True to accept a list of (lat, lon) tuples (distance / bounding-box queries match if any point satisfies the predicate; values are read back as a list of tuples). See doc_values above. |
add_geo3d_field(name, *, stored=True, indexed=True, multi_valued=False, doc_values=True) | 3D ECEF Cartesian point field (x, y, z in metres). Set multi_valued=True to accept a list of (x, y, z) tuples (distance / bounding-box / nearest queries match if any point satisfies the predicate; values are read back as a list of tuples). See Geo3d concepts and doc_values above. |
add_datetime_field(name, *, stored=True, indexed=True, multi_valued=False, doc_values=True) | UTC datetime field. Set multi_valued=True to accept a list of datetime.datetime / str values (range queries match if any instant satisfies the predicate; values are read back as a list[str] of RFC 3339 strings in UTC). See doc_values above. |
add_hnsw_field(name, dimension, *, distance="cosine", m=16, ef_construction=200, quantizer=None, subvector_count=None, rerank_storage=None, embedder=None, pq_codebook_path=None, base_weight=1.0) | HNSW approximate nearest-neighbor vector field. base_weight sets this field’s relative scoring priority when searched alongside other vector fields (Issue #1084); see Vector Search → Weights. |
add_flat_field(name, dimension, *, distance="cosine", embedder=None, base_weight=1.0) | Flat (brute-force) vector field. |
add_ivf_field(name, dimension, *, distance="cosine", n_clusters=100, n_probe=1, embedder=None, base_weight=1.0) | IVF approximate nearest-neighbor vector field. |
add_multi_vector_field(name, dimension, *, distance="cosine", storage="f32", embedder=None) | Multi-vector field holding a variable number of token vectors per document (e.g. ColBERT embeddings), read only by the late-interaction rescore (Issue #1351). It has no ANN index and is not a vector-search target. distance must be "cosine" (default; vectors are L2-normalized when written) or "dot_product", and dimension must be greater than 0; both are checked here and raise ValueError. storage sets the on-disk element kind of each token vector — "f32" (default; exact), "f16" (2x smaller), or "int8" (~4x smaller, lossy quantization); an unrecognized value also raises ValueError. embedder names a token-level embedder ("candle_colbert", see Embedder types) that embeds text values and the text query of a rescore. A document gives the field a list of float lists (see Field value types) or, with embedder, text. The token vectors live only in the vector store: get_documents and search results do not return them. See Multi-Vector Fields. |
Vector quantization & rerank storage (HNSW fields):
quantizer—"scalar_8bit"(default, 4× compression) or"product_quantization"for higher compression. Product quantization requiressubvector_count(must dividedimension).rerank_storage— set to"f32"to write a full-precision*.hnsw.f32sidecar enabling exact Stage-2 rerank; omit to keep the int8-only segment.pq_codebook_path— storage-relative file name of a shared PQ codebook (Issue #631), trained once via thelaurus train pq-codebookCLI command. Only meaningful withquantizer="product_quantization"; commits then encode against the pre-trained codebook instead of re-training k-means per segment. Omit to keep per-segment training.
Every add_*_field method above raises ValueError when name starts with _ (other than _id), and adds nothing. A schema loaded with from_toml / from_toml_file keeps accepting such a field, so a persisted schema still loads; creating a new Index from it raises ValueError. See Field Naming.
Other methods
| Method | Description |
|---|---|
add_embedder(name, config) | Register a named embedder definition. config is a dict with a "type" key (see below), read with the same rules as the schema TOML’s [embedders] table. A config that is not a dict, or has a missing or unknown "type" or a missing required key, raises ValueError (invalid embedder config: ...). |
add_analyzer(name, tokenizer, *, char_filters=None, token_filters=None) | Register a custom analyzer definition. tokenizer is required; char_filters/token_filters are optional lists of dicts. Each dict uses the same {"type": "..."} shape as the schema TOML/JSON format (see below). A name reserved for a built-in analyzer (standard, keyword, english, simple, noop) raises ValueError, and so does creating a new Index from a schema that defines one (e.g. loaded with from_toml). Semantic validity (e.g. a malformed regex) is checked when the schema is used to build an Index, not here. |
analyzer_names() | Return the names of custom analyzers registered via add_analyzer or loaded from TOML. |
Schema.from_toml(toml_str) (static) | Parse a schema from a TOML string, in the same format laurus-cli create index --schema accepts. |
Schema.from_toml_file(path) (static) | Load a schema from a TOML file (path accepts str or os.PathLike). |
to_toml() | Serialize this schema to a TOML string in the same format laurus-cli accepts. |
to_toml_file(path) | Write this schema to a TOML file. |
set_default_fields(fields) | Set default search fields (list of strings). |
set_dynamic_field_policy(policy) | Set how undeclared fields are handled. policy is "strict", "dynamic" (default), or "ignore". See notes below. |
dynamic_field_policy() | Return the current policy as a lowercase string. |
field_names() | Return all field names. |
Dynamic field policy
Controls what happens when a document is ingested with field names that are not declared in the schema:
"strict"— Reject the document."dynamic"(default) — Infer a type for each undeclared field and add it to the schema. Warning: integer fields silently truncate incoming float values (3.14→3). Use"strict"if you need to reject such type mismatches."ignore"— Silently drop the undeclared fields.
See Schema & Fields for the full behaviour matrix.
Embedder types
See Schema Format Reference → Embedders for the canonical description of each type.
"type" | Required keys | Feature flag |
|---|---|---|
"precomputed" | – | (always available) |
"candle_bert" | "model" | embeddings-candle |
"candle_clip" | "model" | embeddings-multimodal |
"openai" | "model" | embeddings-openai |
"candle_colbert" | "model" | embeddings-candle |
"candle_colbert" is a token-level embedder: it produces one vector per token and serves only a multi-vector field (add_multi_vector_field). It also takes the optional keys "revision" (branch, tag, or commit of the model repository), "query_maxlen", and "doc_maxlen"; see Schema Format Reference → Embedders for their defaults.
schema = laurus.Schema()
schema.add_text_field("body")
schema.add_embedder("colbert", {"type": "candle_colbert", "model": "colbert-ir/colbertv2.0"})
schema.add_multi_vector_field("body_colbert", dimension=128, embedder="colbert")
Analyzer components
Used by add_analyzer(name, tokenizer, *, char_filters=None, token_filters=None)
and by the [analyzers.<name>] TOML section. tokenizer is a single dict;
char_filters/token_filters are lists of dicts, applied in list order.
See Schema Format Reference → Analyzers for the canonical description of each component.
Tokenizers (tokenizer, exactly one):
"type" | Required keys | Optional keys |
|---|---|---|
"whitespace" | – | – |
"unicode_word" | – | – |
"regex" | – | "pattern" (default \w+), "gaps" (default false) |
"ngram" | "min_gram", "max_gram" | – |
"lindera" | "mode", "dict" | "user_dict" |
"whole" | – | – |
Char filters (char_filters, applied to raw text before tokenization):
"type" | Required keys | Optional keys |
|---|---|---|
"unicode_normalization" | "form" ("nfc"/"nfd"/"nfkc"/"nfkd") | – |
"pattern_replace" | "pattern", "replacement" | – |
"mapping" | "mapping" (dict of string replacements) | – |
"japanese_iteration_mark" | – | "kanji" (default true), "kana" (default true) |
Token filters (token_filters, applied to the token stream after tokenization):
"type" | Required keys | Optional keys |
|---|---|---|
"lowercase" | – | – |
"stop" | – | "words" (default: English stop words) |
"stem" | – | "stem_type" ("porter"/"simple"/"identity") |
"boost" | "boost" | – |
"limit" | "limit" | – |
"strip" | – | – |
"remove_empty" | – | – |
"flatten_graph" | – | – |
schema = laurus.Schema()
schema.add_analyzer(
"ja_ipadic",
{"type": "lindera", "mode": "normal", "dict": "/var/lib/lindera/ipadic"},
char_filters=[
{"type": "unicode_normalization", "form": "nfkc"},
{"type": "japanese_iteration_mark"},
],
token_filters=[{"type": "lowercase"}],
)
schema.add_text_field("title", analyzer="ja_ipadic")
Distance metrics
| Value | Description |
|---|---|
"cosine" | Cosine similarity (default) |
"euclidean" | Euclidean distance |
"dot_product" | Dot product |
"manhattan" | Manhattan distance |
"angular" | Angular distance |
Query classes
TermQuery
TermQuery(field: str, term: str)
Matches documents containing the exact term in the given field.
PhraseQuery
PhraseQuery(field: str, terms: list[str])
Matches documents containing the terms in order.
FuzzyQuery
FuzzyQuery(field: str, term: str, *, max_edits: int = 2)
Approximate match allowing up to max_edits edit-distance errors. max_edits is keyword-only.
WildcardQuery
WildcardQuery(field: str, pattern: str)
Pattern match. * matches any sequence of characters, ? matches any single character.
NumericRangeQuery
NumericRangeQuery(field: str, *, min: int | float | None = None, max: int | float | None = None)
Matches numeric values in the range [min, max]. Pass None (or omit) for
an open bound. min and max are keyword-only. The numeric type (integer or
float) is inferred from the Python type of min/max.
DateTimeRangeQuery
DateTimeRangeQuery(
field: str, *,
min: str | datetime.datetime | datetime.date | None = None,
max: str | datetime.datetime | datetime.date | None = None,
)
Matches DateTime values in the range [min, max] (both bounds inclusive).
Pass None (or omit) for an open bound; min and max are keyword-only. A
bound is a str literal in any form the query DSL accepts — RFC 3339
("2024-01-01T09:00:00+09:00", normalized to UTC), naive
"YYYY-MM-DDTHH:MM:SS[.fff]" (UTC), or "YYYY-MM-DD" (midnight UTC) — or a
datetime.datetime / datetime.date, converted via isoformat() (a naive
datetime is UTC). A malformed bound raises ValueError at construction.
Usable anywhere a query object is accepted (Index.search, BooleanQuery,
SearchRequest), e.g.
index.search(laurus.DateTimeRangeQuery("created_at", min="2024-01-01", max="2024-12-31")).
GeoDistanceQuery
GeoDistanceQuery.within_radius(
field: str, lat: float, lon: float, distance_m: float,
)
Geo-distance (radius) search. Returns documents whose (lat, lon) coordinate
is within distance_m metres of the given point.
GeoBoundingBoxQuery
GeoBoundingBoxQuery.within_bounding_box(
field: str,
min_lat: float, min_lon: float,
max_lat: float, max_lon: float,
)
Geo bounding-box search. Returns documents whose (lat, lon) coordinate lies
inside the axis-aligned [min_lat, max_lat] × [min_lon, max_lon] rectangle.
Geo3dDistanceQuery
Geo3dDistanceQuery.within_sphere(
field: str, x: float, y: float, z: float, distance_m: float,
)
Sphere search over a 3D ECEF point field. Returns documents whose (x, y, z)
coordinate is within distance_m metres of the centre. See
Geo3d concepts for ECEF theory.
Geo3dBoundingBoxQuery
Geo3dBoundingBoxQuery.within_box(
field: str,
min_x: float, min_y: float, min_z: float,
max_x: float, max_y: float, max_z: float,
)
Axis-aligned 3D bounding-box search. Returns documents whose ECEF point lies
inside [min_x, max_x] × [min_y, max_y] × [min_z, max_z].
Geo3dNearestQuery
Geo3dNearestQuery.k_nearest(
field: str,
x: float, y: float, z: float,
k: int,
*,
initial_radius_m: float | None = None,
max_radius_m: float | None = None,
)
k-nearest-neighbour search over a 3D ECEF point field. Returns the k
documents closest to (x, y, z). The optional initial_radius_m and
max_radius_m tune the iterative-expansion search cone.
BooleanQuery
bq = BooleanQuery()
bq.must(query)
bq.should(query)
bq.must_not(query)
Compound boolean query. Construct with no arguments and add clauses one at a
time via the must / should / must_not methods. Each method accepts any
query object (including a nested BooleanQuery).
must clauses all have to match; must_not clauses must not match.
should clauses contribute to scoring; at least one of them must match if
there are no must clauses.
SpanQuery
# Single term
SpanQuery.term(field: str, term: str)
# Near: terms appearing within `slop` positions of each other
SpanQuery.near(field: str, terms: list[str], *, slop: int = 0, ordered: bool = True)
# Near with nested SpanQuery clauses
SpanQuery.near_spans(field: str, clauses: list[SpanQuery], *, slop: int = 0, ordered: bool = True)
# Containing: big span contains little span
SpanQuery.containing(field: str, big: SpanQuery, little: SpanQuery)
# Within: include span within exclude span at max distance
SpanQuery.within(field: str, include: SpanQuery, exclude: SpanQuery, distance: int)
Positional / proximity span queries. Construct via the static factory
methods. near takes a list of term strings, while near_spans takes a
list of SpanQuery objects for nested expressions. slop and ordered
are keyword-only.
VectorQuery
VectorQuery(field: str, vector: list[float])
Approximate nearest-neighbor search using a pre-computed embedding vector.
VectorTextQuery
VectorTextQuery(field: str, text: str)
Converts text to an embedding at query time and runs vector search. Requires an embedder configured on the index.
SearchRequest
Full-featured search request for advanced control.
class SearchRequest:
def __init__(
self,
*,
query=None,
lexical_query=None,
vector_query=None,
filter_query=None,
fusion=None,
limit: int = 10,
offset: int = 0,
highlight=None,
rescore: LateInteractionRescore | None = None,
) -> None: ...
| Parameter | Description |
|---|---|
query | A DSL string or single query object. Mutually exclusive with lexical_query / vector_query. |
lexical_query | Lexical component for explicit hybrid search. |
vector_query | Vector component for explicit hybrid search. |
filter_query | Lexical filter applied after scoring. |
fusion | Fusion algorithm (RRF or WeightedSum). Defaults to RRF(k=60) when both components are set. |
limit | Maximum number of results (default 10). |
offset | Pagination offset (default 0). |
highlight | Same list-or-dict shape as Index.search’s highlight parameter (Issue #1134). See Highlighting. |
rescore | A LateInteractionRescore applied to the top results of this request, whichever query shape is set (Issue #1351). It takes the place of Index.search’s rescore keyword, which is ignored when a SearchRequest is passed. |
LateInteractionRescore
Late-interaction (ColBERT MaxSim) rescore of the top search results (Issue #1351). Pass it as rescore= to Index.search or SearchRequest.
class LateInteractionRescore:
def __init__(
self,
field: str,
query: str | list[list[float]],
*,
window_size: int | None = None,
) -> None: ...
@property
def window_size(self) -> int: ...
| Parameter | Type | Default | Description |
|---|---|---|---|
field | str | – | A multi-vector field (see add_multi_vector_field). |
query | str | list[list[float]] | – | The query text, embedded by the field’s token-level embedder (a "candle_colbert" one), or the query’s token vectors, computed with the same model as the documents’ token vectors. Integers are accepted as vector elements. |
window_size | int | None | None (100) | How many top first-stage results to rescore, 1..=10,000. Keyword-only. |
The window_size property returns the window the rescore uses (100 when omitted).
The top window_size first-stage results — lexical, vector, or hybrid — are reordered by their MaxSim against the field, highest first, and a rescored result’s score is its MaxSim. Window results without token vectors in the field follow in their first-stage order, and results beyond the window follow last, keeping their first-stage order and score. offset and limit are cut from this rescored ranking. See Vector Search → Late-Interaction Rescore for the ordering, similarity, and cost details.
Errors:
TypeErrorat construction whenqueryis neither astrnor a list of lists, or a token vector holds aboolorstr.ValueError(Invalid argument: rescore: ...) fromIndex.searchfor every other invalid value, checked before the search runs:fieldis not a multi-vector field, the query does not hold 1 to 1,024 vectors of the field’s dimension with finite values,window_sizeis outside1..=10,000, or a text query is blank or the field has no token-level embedder.
import laurus
schema = laurus.Schema()
schema.add_text_field("title")
schema.add_multi_vector_field("tokens", dimension=2, distance="dot_product")
index = laurus.Index(schema=schema)
index.put_document("a", {"title": "rust", "tokens": [[0.1, 0.0]]})
index.put_document("b", {"title": "rust language", "tokens": [[0.9, 0.2]]})
index.commit()
rescore = laurus.LateInteractionRescore("tokens", [[1.0, 0.0], [0.0, 1.0]], window_size=50)
results = index.search("title:rust", rescore=rescore) # "b" (MaxSim 1.1) before "a" (0.1)
# The same rescore inside a SearchRequest (also usable in search_batch)
request = laurus.SearchRequest(query="title:rust", rescore=rescore, limit=5)
results = index.search(request)
With a "candle_colbert" embedder on the field, documents can give the field text and the query can be text too: laurus.LateInteractionRescore("body_colbert", "how do lifetimes work").
SearchResult
Returned by Index.search().
class SearchResult:
id: str # External document identifier
score: float # Relevance score
document: dict | None # Retrieved field values, or None if not stored
highlights: dict[str, list[str]] # Highlighted fragments per requested field
highlights maps each field named in highlight to its highlighted fragments (best first); a field that did not highlight is absent from the dict, and highlights is {} when highlight was not requested. See Highlighting.
With rescore, a rescored result’s score is its MaxSim, while results outside the rescore window keep their first-stage score. See LateInteractionRescore.
Fusion algorithms
RRF
RRF(k: float = 60.0)
Reciprocal Rank Fusion. Merges lexical and vector result lists by rank position. k is a smoothing constant; higher values reduce the influence of top-ranked results.
WeightedSum
WeightedSum(lexical_weight: float = 0.5, vector_weight: float = 0.5)
Normalises both score lists independently, then combines them as lexical_weight * lexical_score + vector_weight * vector_score.
Text analysis
SynonymDictionary
class SynonymDictionary:
def __init__(self) -> None: ...
def add_synonym_group(self, synonyms: list[str]) -> None: ...
WhitespaceTokenizer
class WhitespaceTokenizer:
def __init__(self) -> None: ...
def tokenize(self, text: str) -> list[Token]: ...
SynonymGraphFilter
class SynonymGraphFilter:
def __init__(
self,
dictionary: SynonymDictionary,
keep_original: bool = True,
boost: float = 1.0,
) -> None: ...
def apply(self, tokens: list[Token]) -> list[Token]: ...
Token
class Token:
text: str
position: int
start_offset: int
end_offset: int
boost: float
stopped: bool
position_increment: int
position_length: int
token_type: str | None
start_offset and end_offset are UTF-8 byte offsets into the original text. For non-ASCII text they are not str indices; slice the encoded text instead: text.encode()[tok.start_offset:tok.end_offset].decode().
token_type is one of "alphanum", "num", "cjk", "katakana", "hiragana", "hangul", "punctuation", "whitespace", "synonym", "email", "url" and "other", or None. SynonymGraphFilter.apply keeps each token’s offsets and type, and gives each synonym it inserts the type "synonym" and the offsets of the words it replaces.
Field value types
Python values are automatically converted to Laurus DataValue types:
| Python type | Laurus type | Notes |
|---|---|---|
None | Null | |
bool | Bool | Checked before int |
int | Int64 | |
float | Float64 | |
str | Text | |
bytes | Bytes | |
list[bool] | BoolArray | Multi-valued boolean field; checked before the list[int] rule because bool is a subclass of int. Requires multi_valued=True on the field |
list[bytes] | BytesArray | Multi-valued bytes field (Issue #1176); each element takes the same direct-bytes path as a single bytes value, so MIME is always None. Bytes is never indexed, so this has no query-matching semantics — it only governs the stored shape. Requires multi_valued=True on the field |
list[int] | Int64Array | Multi-valued integer field (a list of bool becomes a BoolArray instead, which a multi-valued Float / Integer field widens to 0/1 in the core); vector fields cast the list to f32. An empty list is an empty Int64Array |
list[float | int] | Float64Array | Multi-valued float field (integers widened); vector fields cast the list to f32 |
list[list[float | int]] | VectorArray | Token vectors of a multi-vector field (Issue #1351), one inner list per token, e.g. {"tokens": [[0.1, 0.2], [0.3, 0.4]]}. Integers are widened; a bool or str element raises TypeError. The vectors’ count and dimension are checked against the field when the document is written (ValueError, e.g. for ragged lists). Not returned by get_documents or search results |
(lat, lon) tuple | Geo | Two float values |
(x, y, z) tuple | Geo3d | Three float values (ECEF Cartesian, metres) |
list[(lat, lon)] | GeoArray | List of (lat, lon) tuples; requires multi_valued=True on the field |
list[(x, y, z)] | GeoEcefArray | List of (x, y, z) tuples; requires multi_valued=True on the field |
datetime.datetime | DateTime | Converted via isoformat() |
list[datetime.datetime | str] (containing at least one datetime) | DateTimeArray | Each element parsed like a single datetime (RFC 3339 / ISO 8601 with offset, or naive YYYY-MM-DDTHH:MM:SS as UTC; objects with isoformat() are converted first); a non-datetime element raises ValueError. Requires multi_valued=True on the field |
list[str] | DateTimeArray or TextArray | A DateTimeArray when every element parses as a datetime (same forms as above), a TextArray otherwise (Issue #1175; read back as a list[str]). Requires multi_valued=True on the field |