Browse State-of-the-Art › Topic Classification
Topic Classification
75 papers with code · 0 benchmarks · 10 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 75 papers with code (186 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
16 Jun 2017 5 repositories listedThis paper intends to develop a so-called active learning process for automatically annotating French language tweets that deal with the image (i.
-
20 May 2021 4 repositories listedWe introduce Korean Language Understanding Evaluation (KLUE) benchmark.
-
29 Apr 2021 3 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedLarge pre-trained language models (LMs) have demonstrated remarkable ability as few-shot learners.
-
23 Oct 2019 3 repositories listedBERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm.
-
21 Feb 2024 2 repositories listedData scarcity in low-resource languages can be addressed with word-to-word translations from labeled task data in high-resource languages using bilingual lexicons.
-
14 Sep 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedDespite the progress we have recorded in the last few years in multilingual natural language processing, evaluation is typically limited to a small set of languages with available datasets which excludes a large number…
-
25 May 2022 2 repositories listedThe ability of generative language models (GLMs) to generate text has improved considerably in the last few years, enabling their use for generative data augmentation.
-
13 Oct 2020 2 repositories listedEven though Variational Autoencoders (VAEs) are widely used for semi-supervised learning, the reason why they work remains unclear.
-
4 Aug 2010 2 repositories listedFrom these correspondences a cross-lingual representation is created that enables the transfer of classification knowledge from the source to the target language.
-
17 May 2025 1 repository listedThis work presents a large-scale human-annotated multi-task benchmark dataset for abusive language detection in Tigrinya social media with joint annotations for three tasks: abusiveness, sentiment, and topic…
-
16 May 2025 1 repository listedThis paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents.
-
2 Apr 2025 1 repository listedAutomatic text classification (ATC) has experienced remarkable advancements in the past decade, best exemplified by recent small and large language models (SLMs and LLMs), leveraged by Transformer architectures.
-
18 Feb 2025 1 repository listedOscar Wilde said, "The difference between literature and journalism is that journalism is unreadable, and literature is not read."
-
29 Nov 2024 1 repository listedTo address this challenge, we propose a teacher-student framework based on large language models (LLMs) for developing multilingual news classification models of reasonable size with no need for manual data annotation.
-
22 Oct 2024 1 repository listedThese results emphasize the significance of URLs in search engine optimization: well-named URLs enable better topic classification, increasing the likelihood of appearing on the first page of search engine results by 4.
-
12 Oct 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedWe hypothesize that LMs perform ICL with irrelevant labels via two sequential processes: an inference function that solves the task, followed by a verbalization function that maps the inferred answer to the label space.
-
26 Sep 2024 1 repository listedContextualized embeddings based on large language models (LLMs) are available for various languages, but their coverage is often limited for lower resourced languages.
-
4 Aug 2024 1 repository listedAs NLP models become increasingly integral to decision-making processes, the need for explainability and interpretability has become paramount.
-
19 Jul 2024 1 repository listedUsing the generated LLM annotations, we explore the finetuning of a specialized smaller classification model, to reduce the computational cost.
-
21 Jun 2024 1 repository listedA simple approach often relies on comparing embeddings of query (text) to those of potential classes.
-
13 Jun 2024 1 repository listedA text classifier is used to ensure that we only include newswire articles, which historically are in the public domain.
-
21 May 2024 1 repository listedThis paper addresses a critical gap in legal analytics by developing and applying a novel taxonomy for topic classification of summary judgment cases in the United Kingdom.
-
16 May 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedIt is often desirable to distill the capabilities of large language models (LLMs) into smaller student models due to compute and memory constraints.
-
1 May 2024 1 repository listedIn this paper, we focus on addressing the issues of topic granularity and hallucinations for better LLM-based topic modelling.
-
29 Apr 2024 1 repository listedThis study investigates the factors influencing the performance of multilingual large language models (MLLMs) across diverse languages.
-
28 Apr 2024 1 repository listedWe conduct a comparative analysis between monolingual and multilingual BERT models, including MahaBERT, IndicBERT, and MuRIL.
-
5 Apr 2024 1 repository listedA common method for ZSC is to fine-tune a language model on a Natural Language Inference (NLI) dataset and then use it to infer the entailment between the input document and the target labels.
-
1 Mar 2024 1 repository listedThis work proposes a novel approach that leverages LLMs for topic classification of column headers using a controlled vocabulary, presenting a practical application of LLMs and Large Context Windows within the Semantic…
-
L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages4 Jan 2024 1 repository listedThis research contributes significantly to expanding the pool of available text classification datasets and also makes it possible to develop topic classification models for Indian regional languages.
-
5 Dec 2023 1 repository listedWith the growing volume of diverse information, the demand for classifying arbitrary topics has become increasingly critical.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections