{"url":"/task/text-categorization","name":"Text Categorization","slug":"text-categorization","description_markdown":"**Text Categorization** is the task of automatically assigning pre-defined categories to documents written in natural languages. Several types of Text Categorization have been studied, each of which deals with different types of documents and categories, such as topic categorization to detect discussed topics (e.g., sports, politics), spam detection, and sentiment classification to determine the sentiment typically in product or movie reviews.\n\n\n<span class=\"description-source\">Source: [Effective Use of Word Order for Text Categorization with Convolutional Neural Networks ](https://arxiv.org/abs/1412.1058)</span>","categories":[{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":247,"papers_with_code":44,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":6,"subtasks":0,"parent_tasks":1},"benchmarks":[],"datasets":[{"url":"/dataset/multilingual-reuters","name":"Multilingual Reuters","full_name":"Multilingual Reuters Collection","num_papers_in_archive":13},{"url":"/dataset/dawt","name":"DAWT","full_name":"Densely Annotated Wikipedia Texts","num_papers_in_archive":3},{"url":"/dataset/italian-crime-news","name":"DICE: a Dataset of Italian Crime Event news","full_name":"from Gazzetta di Modena [2011-2021]","num_papers_in_archive":3},{"url":"/dataset/natcat","name":"NatCat","full_name":"","num_papers_in_archive":1},{"url":"/dataset/mnad","name":"MNAD","full_name":"Moroccan News Articles Dataset","num_papers_in_archive":0},{"url":"/dataset/comentarios-vacuna-vph","name":"Text_VPH","full_name":"","num_papers_in_archive":0}],"subtasks":[],"parent_tasks":[{"url":"/task/text-classification","name":"Text Classification"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":44,"tagged_in_all":247,"items":[{"url":"/paper/on-the-role-of-text-preprocessing-in-neural","title":"On the Role of Text Preprocessing in Neural Network Architectures: An Evaluation Study on Text Categorization and Sentiment Analysis","date":"2017-07-06","arxiv_id":"1707.01780","repositories_listed":3,"syntology":null},{"url":"/paper/a-deeper-look-into-sarcastic-tweets-using","title":"A Deeper Look into Sarcastic Tweets Using Deep Convolutional Neural Networks","date":"2016-10-27","arxiv_id":"1610.08815","repositories_listed":3,"syntology":null},{"url":"/paper/learning-to-few-shot-learn-across-diverse","title":"Learning to Few-Shot Learn Across Diverse Natural Language Classification Tasks","date":"2019-11-10","arxiv_id":"1911.03863","repositories_listed":2,"syntology":null},{"url":"/paper/inverse-category-frequency-based-supervised","title":"Inverse-Category-Frequency based supervised term weighting scheme for text categorization","date":"2010-12-13","arxiv_id":"1012.2609","repositories_listed":2,"syntology":null},{"url":"/paper/latent-dirichlet-allocation","title":"Latent Dirichlet Allocation","date":"2003-01-01","arxiv_id":null,"repositories_listed":2,"syntology":null},{"url":"/paper/a-model-ensemble-approach-with-llm-for","title":"A Model Ensemble Approach with LLM for Chinese Text Classification","date":"2024-03-22","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/harnessing-large-language-models-over","title":"Harnessing Large Language Models Over Transformer Models for Detecting Bengali Depressive Social Media Text: A Comprehensive Study","date":"2024-01-14","arxiv_id":"2401.07310","repositories_listed":1,"syntology":null},{"url":"/paper/beyond-original-research-articles","title":"Beyond original Research Articles Categorization via NLP","date":"2023-09-13","arxiv_id":"2309.07020","repositories_listed":1,"syntology":null},{"url":"/paper/quantum-recurrent-neural-networks-for","title":"Quantum Recurrent Neural Networks for Sequential Learning","date":"2023-02-07","arxiv_id":"2302.03244","repositories_listed":1,"syntology":null},{"url":"/paper/improving-pre-trained-weights-through-meta","title":"Improving Pre-Trained Weights Through Meta-Heuristics Fine-Tuning","date":"2022-12-19","arxiv_id":"2212.09447","repositories_listed":1,"syntology":null},{"url":"/paper/very-large-language-model-as-a-unified","title":"Very Large Language Model as a Unified Methodology of Text Mining","date":"2022-12-19","arxiv_id":"2212.09271","repositories_listed":1,"syntology":null},{"url":"/paper/text-ranking-and-classification-using-data","title":"Text Ranking and Classification using Data Compression","date":"2021-09-23","arxiv_id":"2109.11577","repositories_listed":1,"syntology":null},{"url":"/paper/clustering-word-embeddings-with-self","title":"Clustering Word Embeddings with Self-Organizing Maps. Application on LaRoSeDa -- A Large Romanian Sentiment Data Set","date":"2021-01-11","arxiv_id":"2101.04197","repositories_listed":1,"syntology":null},{"url":"/paper/improving-arabic-text-categorization-using","title":"Improving Arabic Text Categorization Using Transformer Training Diversification","date":"2020-12-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/natcat-weakly-supervised-text-classification","title":"NatCat: Weakly Supervised Text Classification with Naturally Annotated Resources","date":"2020-09-29","arxiv_id":"2009.14335","repositories_listed":1,"syntology":null},{"url":"/paper/item-tagging-for-information-retrieval-a","title":"Item Tagging for Information Retrieval: A Tripartite Graph Neural Network based Approach","date":"2020-08-26","arxiv_id":"2008.11567","repositories_listed":1,"syntology":null},{"url":"/paper/sememnn-a-semantic-matrix-based-memory-neural","title":"SeMemNN: A Semantic Matrix-Based Memory Neural Network for Text Classification","date":"2020-03-04","arxiv_id":"2003.01857","repositories_listed":1,"syntology":null},{"url":"/paper/pyss3-a-python-package-implementing-a-novel","title":"PySS3: A Python package implementing a novel text classifier with visualization tools for Explainable AI","date":"2019-12-19","arxiv_id":"1912.09322","repositories_listed":1,"syntology":null},{"url":"/paper/improving-document-classification-with-multi","title":"Improving Document Classification with Multi-Sense Embeddings","date":"2019-11-18","arxiv_id":"1911.07918","repositories_listed":1,"syntology":null},{"url":"/paper/t-ss3-a-text-classifier-with-dynamic-n-grams","title":"t-SS3: a text classifier with dynamic n-grams for early risk detection over text streams","date":"2019-11-11","arxiv_id":"1911.06147","repositories_listed":1,"syntology":null},{"url":"/paper/ensemble-quantile-classifier","title":"Ensemble Quantile Classifier","date":"2019-10-28","arxiv_id":"1910.12960","repositories_listed":1,"syntology":null},{"url":"/paper/text-categorization-by-learning-predominant","title":"Text Categorization by Learning Predominant Sense of Words as Auxiliary Task","date":"2019-07-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/leap-lstm-enhancing-long-short-term-memory","title":"Leap-LSTM: Enhancing Long Short-Term Memory for Text Categorization","date":"2019-05-28","arxiv_id":"1905.11558","repositories_listed":1,"syntology":null},{"url":"/paper/rep-the-set-neural-networks-for-learning-set","title":"Rep the Set: Neural Networks for Learning Set Representations","date":"2019-04-03","arxiv_id":"1904.01962","repositories_listed":1,"syntology":null},{"url":"/paper/learning-graph-pooling-and-hybrid","title":"Learning Graph Pooling and Hybrid Convolutional Operations for Text Representations","date":"2019-01-21","arxiv_id":"1901.06965","repositories_listed":1,"syntology":null},{"url":"/paper/structure-aware-convolutional-neural-networks","title":"Structure-Aware Convolutional Neural Networks","date":"2018-12-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/hft-cnn-learning-hierarchical-category","title":"HFT-CNN: Learning Hierarchical Category Structure for Multi-label Short Text Categorization","date":"2018-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/revisiting-neural-relation-classification-in","title":"Revisiting neural relation classification in clinical notes with external information","date":"2018-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/using-the-tsetlin-machine-to-learn-human","title":"Using the Tsetlin Machine to Learn Human-Interpretable Rules for High-Accuracy Text Categorization with Medical Applications","date":"2018-09-12","arxiv_id":"1809.04547","repositories_listed":1,"syntology":null},{"url":"/paper/seven-augmenting-word-embeddings-with","title":"SeVeN: Augmenting Word Embeddings with Unsupervised Relation Vectors","date":"2018-08-18","arxiv_id":"1808.06068","repositories_listed":1,"syntology":null}],"syntology_records":0,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}