Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Korean

Litsea supports Korean word segmentation with specialized Hangul character type detection.

Character Types

CodeNamePatternExamples
EParticles/Endings[은는을를의에]은, 는, 을, 를, 의, 에
SNHangul (no 받침)Codepoint arithmetic가, 나, 하, 모
SFHangul (with 받침)Codepoint arithmetic한, 글, 각, 붙
JHangul JamoU+1100–U+11FFIndividual consonants/vowels
GCompatibility JamoU+3130–U+318Fㄱ, ㅏ, ㅎ
HHanjaU+4E00–U+9FFFCJK Ideographs
PPunctuationCJK Symbols + Full-width。, ,
AASCII/Latin[a-zA-Za-zA-Z]A, z
NDigits[0-90-9]0, 5, 5
OOtherFallback@, #, $

Korean Particles (조사)

The “E” type captures six high-frequency grammatical particles:

CharacterRoleName
은/는Topic marker주격 조사
을/를Object marker목적격 조사
Possessive관형격 조사
Locative부사격 조사

These particles frequently appear at word boundaries and are given a distinct type code to improve segmentation accuracy.

Hangul Syllable Structure (받침 Detection)

Korean uses a range arm with a codepoint test in its body for the SN and SF types. This exploits the systematic Unicode Hangul encoding:

  • Hangul Syllables: U+AC00–U+D7AF (11,172 syllables)
  • Each syllable = (initial * 21 + medial) * 28 + final + 0xAC00
  • SN (no 받침): (codepoint - 0xAC00) % 28 == 0
  • SF (with 받침): (codepoint - 0xAC00) % 28 != 0

The 받침 (final consonant) distinction is linguistically significant because it affects how particles attach to words and where boundaries occur.

No WC Features

Korean does not use WC (word + character-type) features. Since most Hangul syllables fall into only two types (SN and SF), WC features would produce low-entropy, noisy combinations that hurt model accuracy.

Space-Preserving Training

Korean is written with spaces between eojeol (word phrases), and those spaces mark most word boundaries. The Korean model is therefore trained on a space-preserving TSV corpus: tokens are tab-separated and each inter-eojeol space is kept as its own token, so the training text contains the space characters of the original sentence and the model can use them as boundary context. Generate the corpus with corpus_udtreebank.sh -s (which reconstructs spacing from the treebank’s SpaceAfter annotations) and extract features with litsea extract --format tsv. At inference no special handling is needed: segment() receives the spaced text as-is and emits each space as its own token.

Because each space is its own single-character token, the character-level labeling (see AdaBoost) marks two separate boundaries around it: the space itself starts a new token (label B), and the character immediately following the space starts the next token (also label B). Both are deterministic given the corpus construction – a space is always exactly one token, and whatever follows it always begins the next token – so the model learns them as a near-trivial rule. Only the second of these (the boundary that starts the following real word) affects the held-out Word F1 score, since pure-whitespace tokens are excluded from scoring (see Evaluation).

Pre-trained Models

korean.model

  • Training corpus: UD Korean-GSD (space-preserving TSV corpus)
  • Training options: --format tsv --tag-free, 30 epochs of Averaged Perceptron training, collapsed to AdaBoost scalar weights, not pruned (3,132 features) – see Training Procedure for the full recipe
  • Word F1 (held-out): 99.91%
  • Boundary F1 (held-out): 99.96%

The model is trained without the 16 tag-dependent feature templates (--tag-free, issue #183): with the space signal available they measured as contributing nothing, and a pointwise model lets segment() skip its sequential scoring pass entirely. See Tag-Free (Pointwise) Models.

Held-out metrics are computed on the original spaced text with space tokens excluded from scoring.

korean_pos.model

  • Algorithm: two-stage segmentation + POS tagging (a binary boundary classifier plus a word-level tagger with a candidate-tag lexicon)
  • Word F1 (held-out): 99.88%
  • Tagged Word F1 (held-out): 93.95%
  • Note: this model is trained on the same space-preserving corpus as korean.model (issue #198), so its Word F1 sits 0.03pt from korean.model’s 99.91%. Until #198 the two-stage pipeline trained on the unspaced word/POS corpus and scored 94.01% on real spaced input; switching protocols gained +5.9pt Word F1 and +10.8pt tagged-word F1. Spaces are re-emitted as their own tokens tagged X, deterministically, via a single-candidate lexicon entry rather than a classifier guess. See English, where the same change was worth over 20pt because English orthography carries far less boundary signal without spaces
  • Details: see Pre-trained Models

Example

echo "한국어 단어 분할 테스트입니다." | litsea segment -l korean ./models/korean.model