Papers › Which Encoding is the Best for Text Classification in Chinese, English, Japanese and Korean?

Which Encoding is the Best for Text Classification in Chinese, English, Japanese and Korean?

8 Aug 2017arXiv:1708.02657archive 2025-07-28

Xiang Zhang, Yann Lecun

This article offers an empirical study on the different ways of encoding Chinese, Japanese, Korean (CJK) and English languages for text classification. Different encoding levels are studied, including UTF-8 bytes, characters, words, romanized characters and romanized words. For all encoding levels, whenever applicable, we provide comparisons with linear models, fastText and convolutional networks. For convolutional networks, we compare between encoding mechanisms using character glyph images, one-hot (or one-of-n) encoding, and embedding. In total there are 473 models, using 14 large-scale text classification datasets in 4 languages including Chinese, English, Japanese and Korean. Some conclusions from these results include that byte-level one-hot encoding based on UTF-8 consistently produces competitive results for convolutional networks, that word-level n-grams linear models are competitive even without perfect word segmentation, and that fastText provides the best result using character-level n-gram encoding but can overfit when the features are overly rich.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Nov05/Genre-Fiction-Classification mentioned on GitHubpytorch report
dbiir/UER-py mentioned on GitHubpytorch report
zhangxiangxiao/glyph mentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

General ClassificationText Classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

fastText

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections