Papers › An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification

An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification

23 Dec 2024arXiv:2412.17361archive 2025-07-28

Andre Rusli, Makoto Shishido

This study investigates the performance of three popular tokenization tools: MeCab, Sudachi, and SentencePiece, when applied as a preprocessing step for sentiment-based text classification of Japanese texts. Using Term Frequency-Inverse Document Frequency (TF-IDF) vectorization, we evaluate two traditional machine learning classifiers: Multinomial Naive Bayes and Logistic Regression. The results reveal that Sudachi produces tokens closely aligned with dictionary definitions, while MeCab and SentencePiece demonstrate faster processing speeds. The combination of SentencePiece, TF-IDF, and Logistic Regression outperforms the other alternatives in terms of classification performance.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Text Classificationregressiontext-classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

BPELogistic RegressionSentencePiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections