--- title: Chinese Compatible description: A simple tokenizer for Chinese, Japanese, and Korean characters canonical: https://www.paradedb.com/docs/reference/tokenizers/available-tokenizers/chinese-compatible --- The Chinese compatible tokenizer is like the [simple](/reference/tokenizers/available-tokenizers/simple) tokenizer -- it lowercases non-CJK characters and splits on any non-alphanumeric character. Additionally, it treats each CJK character as its own token. ```sql CREATE INDEX search_idx ON mock_items USING paradedb (id, (description::pdb.chinese_compatible)) WITH (key_field='id'); ``` To get a feel for this tokenizer, run the following command and replace the text with your own: ```sql SELECT 'Hello world! 你好!'::pdb.chinese_compatible::text[]; ``` ```ini Expected Response text --------------------- {hello,world,你,好} (1 row) ``` ## Convert Between Traditional and Simplified Chinese Use `chinese_convert` to convert Traditional and Simplified Chinese to the same form before tokenization, so queries can match documents written in either form. ```sql CREATE INDEX search_idx ON mock_items USING paradedb (id, (description::pdb.chinese_compatible('chinese_convert=t2s'))) WITH (key_field='id'); ``` The following conversion modes are supported: | Mode | Description | | ------- | --------------------------------------------- | | `t2s` | Traditional to Simplified | | `s2t` | Simplified to Traditional | | `tw2s` | Traditional Taiwan to Simplified | | `tw2sp` | Traditional Taiwan to Simplified, with idioms | | `s2tw` | Simplified to Traditional Taiwan | | `s2twp` | Simplified to Traditional Taiwan, with idioms |