CJK: ICU words and bigrams
Why 清华 finds 清华大学, and how trad↔simp folding lets one query match both scripts.
Getting Chinese, Japanese and Korean right is suo's first promise, ahead of speed. Most search engines segment CJK text into whole words and stop there — which means a half-typed word finds nothing, because the half-typed text is not a word yet. suo does not rely on whole-word segmentation alone.
Two token streams
Every searchable attribute is tokenized twice:
- ICU words, from
Intl.Segmenter's dictionary segmentation (清华大学→清华大学,東京都→東京都). These give precision and drive the exactness tier — an ICU whole-word match ranks above everything else. - Character bigrams and unigrams, the overlapping two-character (and single-character) slices of the same text (
清华大学→清华华大大学). These are the recall floor: they are what makes a half-typed query findable at all.
A query is matched against both streams. An ICU whole-word match wins the exactness tier; a bigram-only match still counts, just lower in the ranking cascade.
Why this matters while typing
Whole-word segmentation only matches a complete word:
清华alone does not match an indexed清华大学, because清华is not itself a segmented word.- A half-typed
稍後閱(still typing toward稍後閱讀) segments as稍後閱, splitting mid-token, and the fragment is dropped.
Bigrams close both gaps: 清华 is a bigram of 清华大学, and 稍後閱 still contains 稍後, 後閱, both indexed as bigrams of 稍後閱讀. So search-as-you-type actually works in CJK, not just on Enter.
Highlighting inside a word
A bigram match can land in the middle of a word, so the highlight mark must not visually break the word apart. A mark that starts or ends inside a CJK run gets no inline padding and no margin (清华 highlighted inside 清华大学 reads as one word with emphasis, not two words with a gap). Padding only applies next to Latin text or whitespace. Snippets and highlights are always built from the original text by character offset — never from FTS highlight(), which inserts spaces into CJK and drops punctuation.
Traditional ↔ simplified folding
Normalization includes a deterministic trad↔simp character fold, so a query in one script matches content in the other — a zh-CN reader searching 讨论 finds a zh-TW page that says 討論, and vice versa. This runs at both index time and query time, before segmentation.
Language detection for a query
- Kana present → Japanese.
- Han characters only → the index's dominant Chinese variant.
- Hangul → Korean.
- Otherwise → the index's configured language.
The ICU canary
Cloudflare updates the V8 engine (and its ICU data) without announcing a segmentation change. Each index hashes the segmentation of a fixed set of canary sentences on startup; if the hash drifts from what is stored, a background re-segmentation starts automatically, with bigram matching as the recall floor in the meantime — so a silent ICU upgrade degrades gracefully instead of silently breaking whole-word matches.
Custom dictionary
Product names, brand terms and other strings that should never be split mid-word can be added to an index's dictionary (an index-time setting) so they are protected before segmentation runs.