Browse State-of-the-Art › Text Categorization
Text Categorization
44 papers with code · 0 benchmarks · 6 datasets archive 2025-07-28
Text Categorization is the task of automatically assigning pre-defined categories to documents written in natural languages. Several types of Text Categorization have been studied, each of which deals with different types of documents and categories, such as topic categorization to detect discussed topics (e.g., sports, politics), spam detection, and sentiment classification to determine the sentiment typically in product or movie reviews.
Source: Effective Use of Word Order for Text Categorization with Convolutional Neural Networks
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 44 papers with code (247 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Jul 2017 3 repositories listedIn this paper we investigate the impact of simple text preprocessing decisions (particularly tokenizing, lemmatizing, lowercasing and multiword grouping) on the performance of a standard neural text classifier.
-
27 Oct 2016 3 repositories listedSarcasm detection is a key task for many natural language processing tasks.
-
10 Nov 2019 2 repositories listedLEOPARD is trained with the state-of-the-art transformer architecture and shows better generalization to tasks not seen at all during training, with as few as 4 examples per label.
-
13 Dec 2010 2 repositories listedTerm weighting schemes often dominate the performance of many classifiers, such as kNN, centroid-based classifier and SVMs.
-
1 Jan 2003 2 repositories listedEach topic is, in turn, modeled as an infinite mixture over an underlying set of topic probabilities.
-
22 Mar 2024 1 repository listedAutomatic medical text categorization can assist doctors in efficiently managing patient information.
-
14 Jan 2024 1 repository listedThe study categorized Reddit and X datasets into "Depressive" and "Non-Depressive" segments, translated into Bengali by native speakers with expertise in mental health, resulting in the creation of the Bengali Social…
-
13 Sep 2023 1 repository listedThis work proposes a novel approach to text categorization -- for unknown categories -- in the context of scientific literature, using Natural Language Processing techniques.
-
7 Feb 2023 1 repository listedQuantum neural network (QNN) is one of the promising directions where the near-term noisy intermediate-scale quantum (NISQ) devices could find advantageous applications against classical resources.
-
19 Dec 2022 1 repository listedMachine Learning algorithms have been extensively researched throughout the last decade, leading to unprecedented advances in a broad range of applications, such as image classification and reconstruction, object…
-
19 Dec 2022 1 repository listedText data mining is the process of deriving essential information from language text.
-
23 Sep 2021 1 repository listedA well-known but rarely used approach to text categorization uses conditional entropy estimates computed using data compression tools.
-
11 Jan 2021 1 repository listedRomanian is one of the understudied languages in computational linguistics, with few resources available for the development of natural language processing tools.
-
1 Dec 2020 1 repository listedAutomatic categorization of short texts, such as news headlines and social media posts, has many applications ranging from content analysis to recommendation systems.
-
29 Sep 2020 1 repository listedWe describe NatCat, a large-scale resource for text classification constructed from three data sources: Wikipedia, Stack Exchange, and Reddit.
-
26 Aug 2020 1 repository listedIn this work, we propose to formulate item tagging as a link prediction problem between item nodes and tag nodes.
-
4 Mar 2020 1 repository listedText categorization is the task of assigning labels to documents written in a natural language, and it has numerous real-world applications including sentiment analysis as well as traditional topic assignment tasks.
-
19 Dec 2019 1 repository listedA recently introduced text classifier, called SS3, has obtained state-of-the-art performance on the CLEF's eRisk tasks.
-
18 Nov 2019 1 repository listedThrough extensive experiments on multiple real-world datasets, we show that SCDV-MS embeddings outperform previous state-of-the-art embeddings on multi-class and multi-label text categorization tasks.
-
11 Nov 2019 1 repository listedSS3 was created to deal with ERD problems naturally since: it supports incremental training and classification over text streams, and it can visually explain its rationale.
-
28 Oct 2019 1 repository listedIt is also shown that the ensemble quantile classifier is Bayes optimal under suitable assumptions with asymmetric Laplace distribution inputs.
-
1 Jul 2019 1 repository listedDistributions of the senses of words are often highly skewed and give a strong influence of the domain of a document.
-
28 May 2019 1 repository listedCompared to previous models which can also skip words, our model achieves better trade-offs between performance and efficiency.
-
3 Apr 2019 1 repository listedIn several domains, data objects can be decomposed into sets of simpler objects.
-
21 Jan 2019 1 repository listedAnother limitation of GCN when used on graph-based text representation tasks is that, GCNs do not consider the order information of nodes in graph.
-
1 Dec 2018 1 repository listedConvolutional neural networks (CNNs) are inherently subject to invariable filters that can only aggregate local inputs with the same topological structures.
-
1 Oct 2018 1 repository listedThe lower the HS level, the less the categorization performance.
-
1 Oct 2018 1 repository listedRecently, segment convolutional neural networks have been proposed for end-to-end relation extraction in the clinical domain, achieving results comparable to or outperforming the approaches with heavy manual feature…
-
12 Sep 2018 1 repository listedThe Tsetlin Machine either performs on par with or outperforms all of the evaluated methods on both the 20 Newsgroups and IMDb datasets, as well as on a non-public clinical dataset.
-
18 Aug 2018 1 repository listedFor example, by examining clusters of relation vectors, we observe that relational similarities can be identified at a more abstract level than with traditional word vector differences.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections