Module Design
The litsea library crate is organized into focused modules, each with a clear responsibility.
Module Dependency Graph
graph TD
language["language.rs<br/>Character classification"]
segmenter["segmenter.rs<br/>Segmentation + POS tagging"]
adaboost["adaboost.rs<br/>AdaBoost (boundaries)"]
perceptron["perceptron.rs<br/>Averaged Perceptron (POS)"]
upos["upos.rs<br/>UPOS tags and labels"]
extractor["extractor.rs<br/>Feature extraction"]
trainer["trainer.rs<br/>Training orchestration"]
two_stage["two_stage.rs<br/>Two-stage model container"]
word_features["word_features.rs (private)<br/>Stage-2 word feature templates"]
packed_model["packed_model.rs (private)<br/>Feature templates + packed AdaBoost tables"]
packed_two_stage["packed_two_stage.rs (private)<br/>Packed two-stage tagging tables"]
model_io["model_io.rs (private)<br/>Model URI loading"]
error["error.rs<br/>LitseaError / Result"]
metrics["metrics.rs<br/>Evaluation metrics (in-sample)"]
evaluation["evaluation.rs<br/>Held-out quality metrics"]
language --> segmenter
upos --> segmenter
adaboost --> segmenter
perceptron --> segmenter
packed_model --> segmenter
packed_two_stage --> segmenter
two_stage --> segmenter
segmenter --> extractor
two_stage --> extractor
word_features --> extractor
evaluation --> extractor
adaboost --> trainer
perceptron --> trainer
two_stage --> trainer
adaboost --> two_stage
perceptron --> two_stage
upos --> two_stage
language --> word_features
word_features --> packed_two_stage
language --> packed_two_stage
perceptron --> packed_two_stage
upos --> packed_two_stage
model_io --> adaboost
model_io --> perceptron
error --> adaboost
error --> perceptron
metrics --> trainer
segmenter --> evaluation
upos --> evaluation
Module Details
language.rs – Language Definitions
Defines the Language enum and character type classification.
Language– Enum with variantsJapanese,Chinese,Korean,English- Implements
FromStr(parses"japanese","ja","chinese","zh","korean","ko","english","en") - Implements
Display(outputs lowercase name) char_type(c: char) -> &'static str– Classifies a character as a table lookup over the numeric type id returned by the privatechar_type_id(), which dispatches to a per-language function (japanese_char_type_id, etc.) implemented as a directmatchon character ranges (allocation-free; no regex). The language-specific functions share apunct_latin_digit()helper for the common"P"/"A"/"N"classes.
- Implements
segmenter.rs – Word Segmentation and POS Tagging
The main user-facing module.
Segmenter– Holds aLanguageand anAdaBoostlearner (fields are private; uselanguage(),learner(),learner_mut()), plus an internal cache for the compiled scoring tables (packed) that backsegment(), and an optional compiled two-stage tagging model (set bywith_two_stage_learner, backingsegment_with_pos(); unlike the cache it is the stage-2 model itself – the raw learner parts are dropped after compilation)new(language)– Create a segmenter with a default (empty) AdaBoost learnerwith_learner(language, learner)– Create a segmenter with a pre-configured AdaBoost learner (e.g. one that has loaded a pre-trained model)with_two_stage_learner(language, learner)– Create a segmenter for two-stage segmentation + POS tagging from aTwoStageLearnersegment(sentence)– Segment text into words, returnsVec<String>segment_into(sentence, buf)– Allocation-free variant (#184): returns token byte ranges borrowed from a reusableSegmentBuffersegment_with_pos(sentence)– Segment and tag, returnsResult<Vec<(String, Upos)>>(PosLearnerNotSetunless a two-stage learner is set)char_type(ch)– Classify a single character into its type codeadd_corpus(corpus)/add_corpus_tsv(corpus)– Add training data (space-separated, or tab-separated/space-preserving; the latter is used for Korean and English, see issue #152)add_corpus_with_writer(corpus, callback)/add_corpus_with_pos_writer(corpus, callback)/add_corpus_tsv_with_writer(corpus, callback)– Process a corpus with a custom callback (the POS-writer variant feeds two-stage stage-1 feature extraction)
adaboost.rs – AdaBoost Algorithm
The binary classifier used for word boundary decisions.
AdaBoostnew(threshold, num_iterations)– Create with training parametersinitialize_features(path)/initialize_instances(path)– Load training datatrain(running)– Run the AdaBoost training looppredict(&attributes)– Predict boundary (+1) or non-boundary (-1)load_model(uri)(async) /load_model_from_path(path)/load_model_from_reader(reader)– Load model weightssave_model(path)– Save model weights to a filemetrics()– Calculate accuracy, precision, and recall (BinaryMetrics)bias()– Get the model’s bias term
perceptron.rs – Averaged Perceptron
The multiclass classifier behind two-stage training (both stages) and the bundled segmentation models’ collapse recipe.
AveragedPerceptronadd_instance(features, label)– Add a training instancetrain(num_epochs, running)– Train with weight averaging (running: &AtomicBool)predict(&features)– Predict the best class labelload_model(uri)(async) /load_model_from_path(path)/load_model_from_reader(reader)– Load model weightssave_model(path)– Save model weightsmetrics()– Macro-averaged evaluation (MulticlassMetrics)
- Weights are stored in a feature → per-class vector layout for fast inference.
upos.rs – Universal POS Tags
Upos– The 17 Universal Dependencies POS tags (NOUN,VERB, …)SegmentLabel– Combined segmentation + POS label per character position (B(Upos)orO), withDisplay/FromStrfor the"B-NOUN"/"O"string form
extractor.rs – Feature Extraction
Extracts features from a corpus for model training.
Extractor– Wraps aSegmenterto process corpus filesnew(language)– Create an extractor for a specific languageextract(corpus_path, features_path)– Read a corpus, write a features fileextract_tsv(corpus_path, features_path)– Same for tab-separated, space-preserving corpora (issue #152, used for Korean and English)extract_two_stage(corpus_path, output_prefix, feature_set)– Extract two-stage training features (issue #147) from a POS-tagged corpus: writes{output_prefix}.stage1(boundary features),.stage2(word-level features), and.lexicon
trainer.rs – Training Orchestration
High-level training workflows.
Trainer– Segmentation model training (AdaBoost)new(threshold, num_iterations, features_path)– Initialize from a features fileload_model(uri)– Optionally load an existing model for incremental training (async)train(running, model_path)– Train and save, returnsBinaryMetrics
PerceptronTrainer– Generic Averaged Perceptron training over opaque string labels (the training step of the bundled segmentation models’ collapse recipe)new(num_epochs, features_path)/load_model(uri)/train(running, model_path)returningMulticlassMetrics
TwoStageTrainer– Two-stage model training (issue #147): trains a stage-1 boundaryAveragedPerceptronand a stage-2 word tagger from the filesExtractor::extract_two_stagewrites, then collapses stage 1 to AdaBoost format and assembles aTwoStageLearnernew(num_epochs, dominance, features_prefix)/train(running, model_path)returningTwoStageMetrics(see Trainer for the full API)
TwoStageMetrics– OneMulticlassMetricsper stage of aTwoStageTrainer::trainrun (stage1,stage2)
two_stage.rs – Two-Stage Model Container
Defines the litsea-two-stage v1 file format (see Model File Format) and the types that hold a two-stage model in memory (issue #147).
TwoStageLearner– Bundles a stage-1 boundaryAdaBoostmodel, a stage-2AveragedPerceptronword tagger, and a candidate-tag lexicon;new()/from_parts(...)/load_model_from_path(path)/save_model(path)mirror the single-learner types’ APITwoStageFeatureSet– Enum selecting the stage-2 word-level template subset (Fast,Balanced,Full)ModelKind– Detect a model file’s format from its first line (AdaBoost, standaloneAveragedPerceptron, orTwoStage); used to give wrong-kind files precise loader errors
error.rs – Error Handling
LitseaError– Error enum (Io,InvalidData,InvalidInput,Unsupported,PosLearnerNotSet, andDownloadwith theremote_modelfeature). Marked#[non_exhaustive], so downstreammatchexpressions need a wildcard armResult<T>– Alias used by every fallible API
metrics.rs – Evaluation Metrics
BinaryMetrics– Accuracy, precision, recall, confusion matrix (AdaBoost)MulticlassMetrics– Accuracy and macro-averaged precision/recall (Averaged Perceptron)
evaluation.rs – Held-Out Evaluation Metrics
While metrics.rs reports in-sample quality (measured on the training data itself, as printed by train), this module computes held-out quality: it compares a Segmenter’s output against a gold corpus using character-offset spans, so predicted and gold tokens can be matched exactly regardless of tokenization differences elsewhere in the sentence.
SegmentationMetrics– Word and boundary precision/recall/F1 for word segmentation, produced byevaluate_segmentation(segmenter, gold)PosMetrics– Wraps aSegmentationMetricsplus tagged-word precision/recall/F1, produced byevaluate_pos(segmenter, gold)(fallible: propagatessegment_with_poserrors)parse_gold_line(line, tsv)/parse_gold_pos_line(line)– Parse a gold corpus line into tokens (plain or POS-tagged); also used by the two-stage extractor and trainer- Backs the CLI’s
litsea evaluatesubcommand
packed_model.rs – Feature Templates and Packed AdaBoost Tables (private)
Internal module holding the declarative feature-template table (TEMPLATES, the single source of truth for all feature consumers), the load-time parser that converts model feature strings into packed integer keys, and PackedModel – the AdaBoost weights compiled into the merged/dense tables read by segment()’s two-pass scorer. Not part of the public API.
packed_two_stage.rs – Packed Two-Stage Tagging Tables (private)
Internal module that compiles a TwoStageLearner’s stage-2 tagger and lexicon into the dense/sparse scoring tables segment_with_pos() reads, mirroring what packed_model.rs does for the AdaBoost learner: a surface map covering the lexicon and dominance-skip tags, sparse per-class rows for the char-valued word-feature templates (from word_features.rs), and dense tables for the type-valued and word-length templates. Not part of the public API.
word_features.rs – Stage-2 Word Feature Templates (private)
Internal module defining the word-level feature templates used by the two-stage tagger’s stage 2 (surface, word length, first/last char and type, context chars/types/bigrams, …). It is the single source of truth for the template set: the training extractor (via extract_two_stage) writes feature strings with write_word_features, and packed_two_stage.rs compiles the same strings back into integer keys with parse_word_feature, pinned against each other by a round-trip test. Not part of the public API.
model_io.rs – Model Loading I/O (private)
Internal module that resolves a model URI (plain path, file://, or http(s):// with the remote_model feature) and returns the raw model bytes. Not part of the public API.
Public Exports
The library’s lib.rs exposes the public modules and re-exports the main types:
#![allow(unused)]
fn main() {
pub mod adaboost;
pub mod error;
pub mod evaluation;
pub mod extractor;
pub mod language;
pub mod metrics;
mod model_io;
mod packed_model;
mod packed_two_stage;
pub mod perceptron;
pub mod segmenter;
pub mod trainer;
pub mod two_stage;
pub mod upos;
mod word_features;
pub use adaboost::AdaBoost;
pub use error::{LitseaError, Result};
pub use evaluation::{PosMetrics, SegmentationMetrics};
pub use extractor::Extractor;
pub use language::{Language, ParseLanguageError};
pub use metrics::{BinaryMetrics, MulticlassMetrics};
pub use perceptron::AveragedPerceptron;
pub use segmenter::{SegmentBuffer, Segmenter};
pub use trainer::{PerceptronTrainer, Trainer, TwoStageMetrics, TwoStageTrainer};
pub use two_stage::{
ModelKind, ParseTwoStageFeatureSetError, TwoStageFeatureSet, TwoStageLearner,
};
pub use upos::{ParseSegmentLabelError, ParseUposError, SegmentLabel, Upos};
pub fn version() -> &'static str { ... }
}