--- title: Jieba description: The most advanced Chinese tokenizer that leverages both a dictionary and statistical models canonical: https://www.paradedb.com/docs/reference/tokenizers/available-tokenizers/jieba --- The Jieba tokenizer is a tokenizer for Chinese text that leverages both a dictionary and statistical models. It is generally considered to be better at identifying ambiguous Chinese word boundaries compared to the [Chinese Lindera](/reference/tokenizers/available-tokenizers/lindera) and [Chinese compatible](/reference/tokenizers/available-tokenizers/chinese-compatible) tokenizers, but the tradeoff is that it is slower. ```sql CREATE INDEX search_idx ON mock_items USING paradedb (id, (description::pdb.jieba)) WITH (key_field='id'); ``` To get a feel for this tokenizer, run the following command and replace the text with your own: ```sql SELECT 'Hello world! 你好!'::pdb.jieba::text[]; ``` ```ini Expected Response text -------------------------------- {hello," ",world,!," ",你好,!} (1 row) ``` ## Search Mode The Jieba tokenizer enables `search_mode` by default, which splits compound words into smaller tokens in addition to keeping the original words: ```sql SELECT '南京市长江大桥'::pdb.jieba::text[]; ``` ```ini Expected Response text ------------------------------------------------ {南京,京市,南京市,长江,大桥,长江大桥} (1 row) ``` To keep only the original words, set `search_mode=false`. This can be useful when tokenizing a search query to avoid matching documents that contain only part of a compound word: ```sql SELECT '南京市长江大桥'::pdb.jieba('search_mode=false')::text[]; ``` ```ini Expected Response text --------------------- {南京市,长江大桥} (1 row) ``` ## Convert Between Traditional and Simplified Chinese Use `chinese_convert` to convert Traditional and Simplified Chinese to the same form before tokenization, so queries can match documents written in either form. ```sql CREATE INDEX search_idx ON mock_items USING paradedb (id, (description::pdb.jieba('chinese_convert=t2s'))) WITH (key_field='id'); ``` The following conversion modes are supported: | Mode | Description | | ------- | --------------------------------------------- | | `t2s` | Traditional to Simplified | | `s2t` | Simplified to Traditional | | `tw2s` | Traditional Taiwan to Simplified | | `tw2sp` | Traditional Taiwan to Simplified, with idioms | | `s2tw` | Simplified to Traditional Taiwan | | `s2twp` | Simplified to Traditional Taiwan, with idioms |