Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Schema Format Reference

The schema file defines the structure of your index — what fields exist, their types, and how they are indexed. Laurus uses TOML format for schema files.

Overview

A schema consists of five top-level elements:

# Policy for fields not declared below. Optional — defaults to "dynamic".
dynamic_field_policy = "dynamic"

# Fields to search by default when a query does not specify a field.
default_fields = ["title", "body"]

# Custom analyzer definitions, referenced by name from Text fields. Optional.
[analyzers.<analyzer_name>]
# ... tokenizer, char_filters, token_filters

# Embedder definitions, referenced by name from vector fields. Optional.
[embedders.<embedder_name>]
# ... type and type-specific options

# Field definitions. Each field has a name and a typed configuration.
[fields.<field_name>.<FieldType>]
# ... type-specific options
  • dynamic_field_policy — How the engine treats fields present in an ingested document but absent from this schema. Accepted values: "strict", "dynamic", "ignore". Defaults to "dynamic". See Dynamic Schema for the full semantics and the warning about silent truncation under "dynamic".
  • default_fields — A list of field names used as default search targets by the Query DSL. Only lexical fields (Text, Integer, Float, etc.) can be default fields. This key is optional and defaults to an empty list.
  • analyzers — A map of names to custom text-analysis pipelines. A Text field uses one by naming it in its analyzer option. Optional. See Analyzers.
  • embedders — A map of names to embedding models. A vector field uses one by naming it in its embedder option. Optional. See Embedders.
  • fields — A map of field names to their typed configuration. Each field must specify exactly one field type.

Field Naming

  • Field names are arbitrary strings (e.g., title, body_vec, created_at).
  • Field names starting with _ are reserved for the engine. The only allow-listed name is _id (managed automatically). Declaring any other _-prefixed field is rejected when the index is created, with Field name '_score' is reserved: names starting with '_' are reserved for system fields (allowed: '_id'); create index creates nothing. An index created before this check keeps opening, and the field stays unusable — a document that sets it is still rejected at ingestion, exactly as it is today.
  • Field names must be unique within a schema.

Field Types

Fields fall into two categories: Lexical (for keyword/full-text search) and Vector (for similarity search). A single field cannot be both.

Lexical Fields

Text

Full-text searchable field. Text is processed by the analysis pipeline (tokenization, normalization, stemming, etc.).

[fields.title.Text]
indexed = true               # Whether to index this field for search
stored = true                # Whether to store the original value for retrieval
multi_valued = false         # Whether to accept arrays of strings (Issue #1175)
position_increment_gap = 100 # Positions skipped between the elements of a multi-valued field
term_vectors = true          # Whether to store term positions (for phrase and span queries)
doc_values = true            # Whether to also copy the value into DocValues (for sorting/faceting)
analyzer = "standard"        # Analyzer for indexing and querying this field
OptionTypeDefaultDescription
indexedbooltrueEnables searching this field
storedbooltrueStores the original value so it can be returned in results
multi_valuedboolfalseAccept arrays of strings; a term query matches if any element contains the term (Lucene-style “any match”); a phrase query never spans two elements unless its slop reaches position_increment_gap
position_increment_gapinteger100Positions skipped between the elements of a multi-valued field (Lucene positionIncrementGap); 0 numbers the elements as if concatenated. Ignored unless multi_valued = true
term_vectorsbooltrueStores term positions, read by phrase and span queries; highlighting always re-tokenizes the stored text and does not use them
doc_valuesbooltrueCopies the value into DocValues, the column-oriented store sorting and faceting/aggregation read from. Takes effect only when stored is also true — see Common option: doc_values below
analyzerstring or table(omit)Analyzer used both when indexing and when parsing queries against this field. A string names a built-in analyzer ("standard", "english", "keyword", "simple", "noop") or an entry in [analyzers.*]. A table selects a parameterized built-in preset; today only { language = "japanese", mode = "normal", dict = "<path>" } (see Text Analysis). When omitted, "standard" is used

The interactive generator (laurus create schema, see Generating a Schema) asks whether a Text field is multi-valued and, when it is, for its position increment gap.

Integer

64-bit signed integer field. Supports range queries and exact match.

[fields.year.Integer]
indexed = true
stored = true
multi_valued = false
doc_values = true
OptionTypeDefaultDescription
indexedbooltrueEnables range and exact-match queries
storedbooltrueStores the original value
multi_valuedboolfalseAccept arrays of integers; range queries match if any value satisfies the predicate (Lucene-style “any match” with constant scoring)
doc_valuesbooltrueSee Common option: doc_values below

Float

64-bit floating point field. Supports range queries.

[fields.rating.Float]
indexed = true
stored = true
multi_valued = false
doc_values = true
OptionTypeDefaultDescription
indexedbooltrueEnables range queries
storedbooltrueStores the original value
multi_valuedboolfalseAccept arrays of floats; range queries match if any value satisfies the predicate (Lucene-style “any match” with constant scoring)
doc_valuesbooltrueSee Common option: doc_values below

Boolean

Boolean field (true / false).

[fields.published.Boolean]
indexed = true
stored = true
multi_valued = false
doc_values = true
OptionTypeDefaultDescription
indexedbooltrueEnables filtering by boolean value
storedbooltrueStores the original value
multi_valuedboolfalseAccept arrays of booleans; a term query (flags:true) matches if any element equals the queried value (Lucene-style “any match”); repeated elements raise the term frequency, not the hit count
doc_valuesbooltrueSee Common option: doc_values below

DateTime

UTC timestamp field. Supports range queries.

[fields.created_at.DateTime]
indexed = true
stored = true
multi_valued = false
doc_values = true
OptionTypeDefaultDescription
indexedbooltrueEnables range queries on date/time
storedbooltrueStores the original value
multi_valuedboolfalseAccept arrays of instants; range queries match if any instant satisfies the predicate (Lucene-style “any match”)
doc_valuesbooltrueSee Common option: doc_values below

Geo

Geographic point field (latitude/longitude). Supports radius and bounding box queries.

[fields.location.Geo]
indexed = true
stored = true
multi_valued = false
doc_values = true
OptionTypeDefaultDescription
indexedbooltrueEnables geo queries (radius, bounding box)
storedbooltrueStores the original value
multi_valuedboolfalseAccept arrays of points; distance / bounding-box queries match if any point satisfies the predicate (Lucene-style “any match”), scoring the document by its closest point
doc_valuesbooltrueSee Common option: doc_values below

Geo3d

3D Earth-Centered Earth-Fixed (ECEF) Cartesian point field (x / y / z in meters). Supports the geo3d_distance (sphere), geo3d_bbox (3D AABB), and geo3d_nearest (k-NN) queries. See 3D Geographic Search (ECEF) for the coordinate system and the wgs84_to_ecef / ecef_to_wgs84 conversion utilities.

[fields.position.Geo3d]
indexed = true
stored = true
multi_valued = false
doc_values = true
OptionTypeDefaultDescription
indexedbooltrueEnables 3D geo queries (geo3d_distance, geo3d_bbox, geo3d_nearest)
storedbooltrueStores the original (x, y, z) value
multi_valuedboolfalseAccept arrays of points; geo3d_distance / geo3d_bbox / geo3d_nearest queries match if any point satisfies the predicate (Lucene-style “any match”), scoring the document by its closest point
doc_valuesbooltrueSee Common option: doc_values below

Bytes

Raw binary data field. Not indexed — stored only.

[fields.thumbnail.Bytes]
stored = true
multi_valued = false
OptionTypeDefaultDescription
storedbooltrueStores the binary data
multi_valuedboolfalseAccept arrays of byte strings; a Bytes field is never indexed, so unlike every other multi_valued option this has no “any match” query semantics — it only governs the stored shape and ingestion arity

BytesOption has no doc_values setting: a Bytes value is never written to DocValues regardless, since neither sorting nor faceting can do anything with it.

Common option: doc_values

Every lexical field option above except BytesOption carries a doc_values option, controlling whether the value is also copied into DocValues — the column-oriented store sorting and faceting/aggregation read from. The effective rule: a DocValues column is written only when stored and doc_values are both true. Setting doc_values: false with stored: false is silently ignored (not an error). Turning doc_values off for a field that is never sorted or faceted on shrinks its segment footprint, since the value is then written once (to the stored document) instead of twice; the field remains fully searchable and retrievable either way — sorting and faceting on it simply fall back to the stored document.

Vector Fields

Vector fields are indexed for approximate nearest neighbor (ANN) search. They require a dimension (the length of each vector) and a distance metric.

Hnsw

Hierarchical Navigable Small World graph index. Best for most use cases — offers a good balance of speed and recall.

[fields.body_vec.Hnsw]
dimension = 384
distance = "Cosine"
m = 16
ef_construction = 200
base_weight = 1.0
OptionTypeDefaultDescription
dimensioninteger128Vector dimensionality (must match your embedding model)
distancestring"Cosine"Distance metric (see Distance Metrics)
minteger16Max bi-directional connections per node. Higher = better recall, more memory
ef_constructioninteger200Search width during index construction. Higher = better quality, slower build
base_weightfloat1.0Relative priority vs. other vector fields searched together; no effect on lexical-vs-vector fusion balance (see Vector Search → Weights)
quantizerobject"Scalar8Bit"Quantization method (see Quantization). Mandatory; default keeps the int8 format introduced in Issue #481 Stage 1.
rerank_storagestring(omit)Optional Stage 2 rerank sidecar (see Rerank Storage). "F32" enables a per-field f32 sidecar so search can rescore int8 candidates against the original vectors. Omit to keep Stage 1 int8-only behavior.
pq_codebook_pathstring(omit)Storage-relative file name of a shared PQ codebook (Issue #631); only meaningful with a ProductQuantization quantizer. Train it with laurus train pq-codebook; commits then encode against it instead of re-training k-means per segment. When set but not yet trained, commits fail loudly (no silent fallback). Omit to train per segment.
embedderstring(omit)Name of an entry in [embedders.*]. Text (or image) values given for this field are then embedded with that model, both when indexing and when searching. Omit to supply precomputed vectors only

Tuning guidelines:

  • m: 12–48 is typical. Use higher values for higher-dimensional vectors.
  • ef_construction: 100–500. Higher values produce a better graph but increase build time.
  • dimension: Must exactly match the output dimension of your embedding model (e.g., 384 for all-MiniLM-L6-v2, 768 for BERT-base, 1536 for text-embedding-3-small).

Flat

Brute-force linear scan index. Provides exact results with no approximation. Best for small datasets (< 10,000 vectors).

[fields.embedding.Flat]
dimension = 384
distance = "Cosine"
base_weight = 1.0
OptionTypeDefaultDescription
dimensioninteger128Vector dimensionality
distancestring"Cosine"Distance metric (see Distance Metrics)
base_weightfloat1.0Relative priority vs. other vector fields searched together; no effect on lexical-vs-vector fusion balance (see Vector Search → Weights)
quantizerobject"Scalar8Bit"Quantization method (see Quantization). Mandatory; default keeps the int8 format introduced in Issue #481 Stage 1.
rerank_storagestring(omit)Optional Stage 2 rerank sidecar (see Rerank Storage); supported by all three vector index types since #932. "F32" enables the per-field f32 sidecar so search can rescore int8 candidates against the original vectors.
embedderstring(omit)Name of an entry in [embedders.*]; see Hnsw

Ivf

Inverted File Index. Clusters vectors and searches only a subset of clusters. Suitable for very large datasets.

[fields.embedding.Ivf]
dimension = 384
distance = "Cosine"
n_clusters = 100
n_probe = 1
base_weight = 1.0
OptionTypeDefaultDescription
dimensioninteger(required)Vector dimensionality
distancestring"Cosine"Distance metric (see Distance Metrics)
n_clustersinteger100Number of clusters. More clusters = finer partitioning
n_probeinteger1Number of clusters to search at query time. Higher = better recall, slower
base_weightfloat1.0Relative priority vs. other vector fields searched together; no effect on lexical-vs-vector fusion balance (see Vector Search → Weights)
quantizerobject"Scalar8Bit"Quantization method (see Quantization). Mandatory; default keeps the int8 format introduced in Issue #481 Stage 1.
rerank_storagestring(omit)Optional Stage 2 rerank sidecar (see Rerank Storage); supported by all three vector index types since #932. "F32" enables the per-field f32 sidecar so search can rescore int8 candidates against the original vectors.
embedderstring(omit)Name of an entry in [embedders.*]; see Hnsw

Note: Unlike Hnsw and Flat, the dimension field in Ivf is required and has no default value.

Tuning guidelines:

  • n_clusters: A common heuristic is sqrt(N) where N is the total number of vectors.
  • n_probe: Start with 1 and increase until recall is acceptable. Typical range is 1–20.

MultiVector

Every token vector of a document (for example ColBERT-style per-token embeddings), kept for late-interaction rescoring. It has no ANN index and is not a vector-search target; see Multi-Vector Fields.

[fields.body_colbert.MultiVector]
dimension = 128
distance = "Cosine"
storage = "Int8"  # optional; omit for F32 (exact, default)
OptionTypeDefaultDescription
dimensioninteger128Dimensionality of every token vector
distancestring"Cosine"Token similarity: "Cosine" (vectors are L2-normalized when written) or "DotProduct". Other metrics are rejected
storagestring"F32"On-disk element kind of every token vector: "F32" (exact, 4 bytes/element), "F16" (2 bytes/element, ~2⁻¹¹ relative error/element), or "Int8" (~1 byte/element plus a small per-vector scale, dimension + 2 bytes/row total). See MultiVector Storage. Changing it on an existing field requires a reindex
embedderstring–Name of a token-level embedder (type = "candle_colbert") declared under [embedders]. When set, documents may give the field text, which is embedded into token vectors (see Embedders)

A document’s value is an array of equal-length numeric arrays, between 1 and 8,192 of them, or text when the field has an embedder. The field is not stored in the document store.

Distance Metrics

The distance option for vector fields accepts the following values:

ValueDescriptionUse When
"Cosine"Cosine distance (1 - cosine similarity). Default.Normalized text/image embeddings
"Euclidean"L2 (Euclidean) distanceSpatial data, non-normalized vectors
"Manhattan"L1 (Manhattan) distanceSparse feature vectors
"DotProduct"Dot product (higher = more similar)Pre-normalized vectors where magnitude matters
"Angular"Angular distanceSimilar to cosine, but based on angle

For most embedding models (BERT, Sentence Transformers, OpenAI, etc.), "Cosine" is the correct choice.

Quantization

Vector fields are stored on disk as 8-bit scalar-quantized integers (Issue #481 Stage 1). Quantization is mandatory; the previous “no quantization” mode no longer exists. The quantizer option defaults to Scalar8Bit and can be omitted from TOML.

Scalar 8-bit (default)

Per-segment global affine quantization to u8. Compresses each f32 component to a single byte (~4x memory reduction) with negligible recall loss in practice.

[fields.embedding.Hnsw]
dimension = 384
distance = "Cosine"
# quantizer = "Scalar8Bit"  # implicit default; can be omitted

Product Quantization (HNSW-only)

Issue #481 Stage 3. Stores each vector as subvector_count one-byte centroid indexes against a codebook of 256 centroids per sub-vector (~16-64x compression). Supported by the HNSW index; Flat / IVF reject it at write time. Usually paired with Rerank Storage to recover recall.

[fields.embedding.Hnsw]
dimension = 384
distance = "Cosine"
# Optional (Issue #631): train the codebook once with
# `laurus train pq-codebook` and share it across segments instead of
# re-training k-means on every commit and merge.
pq_codebook_path = "embedding.pqcb"

[fields.embedding.Hnsw.quantizer.ProductQuantization]
subvector_count = 48
OptionTypeDescription
subvector_countintegerNumber of subvectors. Must evenly divide dimension.

By default the codebook is trained per segment (segments with fewer than 256 vectors fall back to Scalar8Bit). With pq_codebook_path set, segments encode against the shared pre-trained codebook instead: commits get dramatically faster, and even tiny per-commit segments stay on PQ — but a commit before the codebook has been trained fails with an error naming the laurus train pq-codebook command to run (never a silent fallback to per-segment training). See the train command for the training workflow.

Breaking change (Issue #481 Stage 1): schemas that explicitly set quantizer to a “none” value are no longer valid. Existing vector indexes built with a pre-Stage-1 laurus build cannot be read; rebuild from source data after upgrading.

Rerank Storage

Optional Stage 2 sidecar (Issue #481) that keeps the original full-precision vectors alongside the int8 segment so the searcher can do a wide candidate fetch over int8 (cheap) and then rescore the top top_k * rerank_factor candidates against the exact f32 values (accurate). Supported by all three vector index types — HNSW, Flat, and IVF (#932); on Flat/IVF the rescoring applies to field-routed queries.

The sidecar is configured per field with rerank_storage:

[fields.embedding.Hnsw]
dimension = 384
distance = "Cosine"
rerank_storage = "F32"  # opt-in; omit for Stage 1 int8-only behavior
ValueOn-disk overheadDescription
"F32"+4 bytes/dim per vectorIEEE-754 single-precision sidecar (Lucene 99 / FAISS convention).

When omitted, no sidecar is written and the field stays on the Stage 1 int8-only search path. Queries that pass rerank_factor against a field without rerank_storage silently fall back to Stage 1 ranking — the searcher cannot recover f32 information that was discarded at index time.

Scope: Stage 2 lands HNSW only. Flat / IVF accept the field for schema symmetry but currently neither emit nor consume the sidecar.

Analyzers

An [analyzers.<name>] table defines a custom text-analysis pipeline. A Text field uses it by naming it in its analyzer option. Define one when no built-in analyzer fits — for example, to add stemming, or to analyze Japanese text without the stop filter of the japanese preset. See Text Analysis for how the pipeline works.

[analyzers.<name>]
char_filters = [{ type = "...", ... }, ...]   # optional
tokenizer = { type = "...", ... }             # required
token_filters = [{ type = "...", ... }, ...]  # optional
KeyTypeDefaultDescription
tokenizertable(required)Splits the text into tokens. Exactly one
char_filtersarray of tables[]Applied to the raw text before tokenization, in array order
token_filtersarray of tables[]Applied to the token stream after tokenization, in array order

Each component is a table whose type key selects the component; its other keys configure it. Inline tables ({ type = "lowercase" }) are the usual TOML spelling. The same shape is used in JSON schemas ({"type": "lowercase"}) and by every binding’s addAnalyzer / add_analyzer.

Tokenizers

typeRequired keysOptional keysDescription
"whitespace"––Splits on whitespace
"unicode_word"––Splits on Unicode word boundaries
"regex"–pattern (default \w+), gaps (default false)Emits each match of pattern as a token. With gaps = true, pattern matches the separators between tokens instead
"ngram"min_gram, max_gram–Emits every n-gram from min_gram to max_gram characters long
"lindera"mode, dictuser_dictMorphological analysis with Lindera. mode is "normal" or "decompose". dict is the path to a Lindera dictionary directory and user_dict the path to a user dictionary; laurus does not embed a dictionary, so dict must exist on disk
"whole"––Emits the whole input as a single token

Char filters

typeRequired keysOptional keysDescription
"unicode_normalization"form ("nfc" / "nfd" / "nfkc" / "nfkd")–Applies Unicode normalization
"pattern_replace"pattern, replacement–Replaces each match of the regular expression pattern with replacement
"mapping"mapping (a table of string replacements)–Replaces each key of mapping with its value
"japanese_iteration_mark"–kanji (default true), kana (default true)Expands Japanese iteration marks (踊り字)

Token filters

typeRequired keysOptional keysDescription
"lowercase"––Lowercases each token
"stop"–words (default: English stop words)Removes stop words
"stem"–stem_type ("porter" (default) / "simple" / "identity")Reduces each token to its stem
"boost"boost–Multiplies each token’s boost by boost
"limit"limit–Keeps at most limit tokens
"strip"––Trims leading and trailing whitespace from each token
"remove_empty"––Removes empty tokens
"flatten_graph"––Flattens a token graph into a linear stream. Indexing already does this; because the analyzer also parses queries, adding it makes quoted multi-word synonyms inexact at query time

Referencing an analyzer

A Text field names an analyzer in its analyzer option:

[fields.body.Text]
analyzer = "english_stemmed"

The name is resolved in this order:

  1. An analyzer registered at runtime through a binding (for example the WASM binding’s addAnalyzer)
  2. A built-in analyzer: standard, keyword, english, simple, or noop
  3. An entry in [analyzers.*]

Because built-ins are checked first, the names standard, keyword, english, simple and noop are reserved: an [analyzers.*] entry under one of them could never be used, so it is rejected. japanese is not reserved, because its built-in needs a dictionary and is selected with a table, so [analyzers.japanese] is used as defined.

Errors surface at three points:

  • An unknown type or a missing required key is rejected when the schema is parsed; create index creates nothing.
  • An entry named after a built-in is rejected when the index is created, with Analyzer name 'standard' is reserved for a built-in analyzer; choose another name; create index creates nothing. Every binding’s addAnalyzer / add_analyzer (WASM: addAnalyzerDefinition) raises the same error. An index created before this check keeps opening and the entry stays unused; opening it emits a warning through the log crate, which laurus-server prints in its log.
  • An invalid value (a malformed regular expression, an unknown form or stem_type, a missing Lindera dictionary) or an analyzer name that resolves to nothing is rejected when the index is built, with an error such as Failed to resolve analyzer for field 'body': ...; create index creates nothing — schema.toml and store/ are rolled back to whatever state (if any) existed before the call.

Example: English text with stemming

default_fields = ["title", "body"]

[analyzers.english_stemmed]
char_filters = [{ type = "unicode_normalization", form = "nfkc" }]
tokenizer = { type = "unicode_word" }
token_filters = [
    { type = "lowercase" },
    { type = "stop" },
    { type = "stem", stem_type = "porter" },
]

[fields.title.Text]
analyzer = "english_stemmed"

[fields.body.Text]
analyzer = "english_stemmed"

[fields.tag.Text]
analyzer = "keyword"

A document whose body is "Dogs are RUNNING in the park." then matches body:dog (NFKC normalization, lowercasing, and stemming) and body:run, while body:the matches nothing (stop words are removed). The tag field keeps the built-in keyword analyzer, so it matches only its exact value.

Example: Japanese text with Lindera

This definition, taken from examples/aozora/schema.toml, differs from the { language = "japanese" } preset in that it has no stop filter, so particles such as の and は stay in the index:

[analyzers.ja_ipadic]
tokenizer = { type = "lindera", mode = "normal", dict = "/var/lib/lindera/ipadic" }
char_filters = [
    { type = "unicode_normalization", form = "nfkc" },
    { type = "japanese_iteration_mark", kanji = true, kana = true },
]
token_filters = [{ type = "lowercase" }]

[fields.title.Text]
analyzer = "ja_ipadic"

dict must point to an unpacked Lindera dictionary (typically IPADIC). If it does not exist, create index fails with Failed to load dictionary: ... Dictionary path does not exist.

Embedders

An [embedders.<name>] table declares an embedding model. A vector field (Hnsw, Flat, Ivf, or MultiVector) uses it by naming it in its embedder option. Text (or, for CLIP, image) values given for that field are then converted to vectors with the model, both when documents are indexed and when a query targets the field. Several fields can share one embedder. See Embeddings for how each model works and how to choose one.

[embedders.<name>]
type = "..."   # required
model = "..."  # required for every type except "precomputed"
typeRequired keysFeature flagDescription
"precomputed"–(always available)Performs no embedding; documents supply the vectors directly
"candle_bert"modelembeddings-candleLocal text embedding with a BERT-family model from Hugging Face Hub, such as "sentence-transformers/all-MiniLM-L6-v2"
"candle_clip"modelembeddings-multimodalLocal text and image embedding with a CLIP model from Hugging Face Hub, such as "openai/clip-vit-base-patch32"
"openai"modelembeddings-openaiText embedding through the OpenAI API, such as "text-embedding-3-small". The API key is read from the OPENAI_API_KEY environment variable when the engine starts and is never stored in the schema
"candle_colbert"modelembeddings-candleLocal token-level embedding with a BERT-based ColBERT checkpoint, such as "colbert-ir/colbertv2.0" or "answerdotai/answerai-colbert-small-v1". Only a MultiVector field can use it

"candle_colbert" also takes these optional keys:

KeyTypeDefaultDescription
revisionstringdefault branchBranch, tag or commit of the model repository. Pin a commit: documents not yet committed are embedded again from the write-ahead log on recovery, and a changed model would produce different vectors
query_maxlenintegerfrom the checkpoint (32)Number of tokens every query is padded or truncated to
doc_maxlenintegerfrom the checkpoint (180 for colbertv2.0, 300 for answerai-colbert-small-v1)Maximum number of tokens of a document

Hugging Face models are downloaded on first use. candle_bert and candle_clip cache them under $HF_HOME (default ~/.cache/huggingface); candle_colbert uses the Hugging Face default cache, $HF_HOME/hub (default ~/.cache/huggingface/hub), which Python tools share. The vector field’s dimension must equal the model’s output dimension; for candle_colbert this is checked when the engine starts.

An embedder must fit the field that names it: a MultiVector field accepts only "candle_colbert" or "precomputed", and an Hnsw, Flat or Ivf field does not accept "candle_colbert". Any other combination is rejected by create index, add field and update field.

Note: The prebuilt release binaries are built with --features embeddings-all. A laurus binary installed with cargo install laurus-cli or built from source enables none of the embedding features unless you pass them (for example cargo install laurus-cli --features embeddings-candle), so only "precomputed" works. See Installation and Feature Flags. A schema that names a type whose feature is missing still parses, but create index fails:

Error: Not implemented: candle_bert embedder requires the 'embeddings-candle' feature to be enabled

A vector field’s embedder must name an [embedders.*] entry. An undeclared name is rejected by create index, add field, and update field, and an existing index whose schema.toml holds one fails to open:

Error: Invalid argument: Unknown embedder 'missing' for field 'vec': not defined in schema.embedders

To open such an index, edit its schema.toml. Declaring the name with type = "precomputed" keeps the field working as it did, with documents supplying its vectors. Deleting the field’s embedder line does the same.

Example: one embedder shared by two fields

[embedders.text_embedder]
type = "candle_bert"
model = "sentence-transformers/all-MiniLM-L6-v2"

[fields.title_vec.Hnsw]
dimension = 384
distance = "Cosine"
embedder = "text_embedder"

[fields.body_vec.Hnsw]
dimension = 384
distance = "Cosine"
embedder = "text_embedder"

Example: ColBERT token vectors for late-interaction rescoring

[embedders.colbert]
type = "candle_colbert"
model = "answerdotai/answerai-colbert-small-v1"
revision = "934fa8bb4ce2284f4c2baa232d81aca4d076fa5e"

[fields.body.Text]
indexed = true
stored = true

[fields.body_colbert.MultiVector]
dimension = 96
distance = "Cosine"
embedder = "colbert"

Documents then give body_colbert the same text as body, and a late-interaction rescore can take its query as text.

Complete Examples

Full-text search only

A simple blog post index with lexical search:

default_fields = ["title", "body"]

[fields.title.Text]
indexed = true
stored = true
term_vectors = true

[fields.body.Text]
indexed = true
stored = true
term_vectors = true

[fields.category.Text]
indexed = true
stored = true
term_vectors = false

[fields.published_at.DateTime]
indexed = true
stored = true

Vector search only

A vector-only index for semantic similarity:

[fields.embedding.Hnsw]
dimension = 768
distance = "Cosine"
m = 16
ef_construction = 200

Hybrid search (lexical + vector)

Combine lexical and vector search for best-of-both-worlds retrieval:

default_fields = ["title", "body"]

[fields.title.Text]
indexed = true
stored = true
term_vectors = true

[fields.body.Text]
indexed = true
stored = true
term_vectors = true

[fields.category.Text]
indexed = true
stored = true
term_vectors = false

[fields.body_vec.Hnsw]
dimension = 384
distance = "Cosine"
m = 16
ef_construction = 200

Tip: A single field cannot be both lexical and vector. Use separate fields (e.g., body for text, body_vec for embedding) and map them both to the same source content.

E-commerce product index

A more complex schema with mixed field types:

default_fields = ["name", "description"]

[fields.name.Text]
indexed = true
stored = true
term_vectors = true

[fields.description.Text]
indexed = true
stored = true
term_vectors = true

[fields.price.Float]
indexed = true
stored = true

[fields.in_stock.Boolean]
indexed = true
stored = true

[fields.created_at.DateTime]
indexed = true
stored = true

[fields.location.Geo]
indexed = true
stored = true

[fields.description_vec.Hnsw]
dimension = 384
distance = "Cosine"

Custom analysis and automatic embedding

A hybrid index whose Text fields use a custom analyzer and whose vector field embeds its text with a local model. It needs a laurus binary with the embeddings-candle feature (see Embedders):

default_fields = ["title", "body"]

[analyzers.english_stemmed]
tokenizer = { type = "unicode_word" }
token_filters = [
    { type = "lowercase" },
    { type = "stop" },
    { type = "stem" },
]

[embedders.text_embedder]
type = "candle_bert"
model = "sentence-transformers/all-MiniLM-L6-v2"

[fields.title.Text]
analyzer = "english_stemmed"

[fields.body.Text]
analyzer = "english_stemmed"

[fields.body_vec.Hnsw]
dimension = 384
distance = "Cosine"
embedder = "text_embedder"

Generating a Schema

You can generate a schema TOML file interactively using the CLI:

laurus create schema
laurus create schema --output my_schema.toml

See create schema for details.

Using a Schema

Once you have a schema file, create an index from it:

laurus create index --schema schema.toml

Or load it programmatically in Rust:

#![allow(unused)]
fn main() {
use laurus::Schema;

let toml_str = std::fs::read_to_string("schema.toml")?;
let schema: Schema = toml::from_str(&toml_str)?;
}