Browse State-of-the-Art › News Classification
News Classification
33 papers with code · 4 benchmarks · 12 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| N15News (9 rows) | Multimodal(ViT+BERT, Input: Image + Body) | N24News: A New Dataset for Multimodal News Classification | code | — | Compare |
| Reddit Ideological and Extreme Bias Dataset (7 rows) | SVM | RICo: Reddit ideological communities | code | — | Compare |
| Soham News Article Classification (3 rows) | xlmindic-base-uniscript | Does Transliteration Help Multilingual Language Modeling? | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| BBC Hindi News Article Classification (2 rows) | xlmindic-base-uniscript | Does Transliteration Help Multilingual Language Modeling? | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 33 papers with code (72 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
19 May 2021 6 repositories listedThe proliferation of fake news, i.
-
1 Aug 2021 4 repositories listed
-
12 Feb 2024 2 repositories listedWe ask the following question in this study: are 01 loss sign activation neural networks hard to deceive with a popular black box text adversarial attack program called TextFooler?
-
14 Mar 2022 2 repositories listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)Many recent deep learning-based solutions have widely adopted the attention-based mechanism in various tasks of the NLP discipline.
-
20 Oct 2021 2 repositories listedIncreasing amounts of freely available data both in textual and relational form offers exploration of richer document representations, potentially improving the model performance and robustness.
-
18 Jul 2025 1 repository listedTo address these challenges, this study proposes a novel deep clustering framework that comprising GCN, Autoencoder (AE), and Graph Transformer, termed the Tri-Learn Graph Fusion Network (Tri-GFN).
-
29 Nov 2024 1 repository listedTo address this challenge, we propose a teacher-student framework based on large language models (LLMs) for developing multilingual news classification models of reasonable size with no need for manual data annotation.
-
26 Nov 2024 1 repository listedNatural Language Processing (NLP) for low-resource languages presents significant challenges, particularly due to the scarcity of high-quality annotated data and linguistic resources.
-
12 Aug 2024 1 repository listedIn total, our combined dataset of 1.
-
5 Jun 2024 1 repository listedPrevious work addressed label curation by using ideological subreddits (r/Liberal and r/Conservative for Liberal and Conservative classes) to label the articles shared on those subreddits according to their prescribed…
-
17 Mar 2024 1 repository listedThis paper presents a comprehensive examination of the impact of tokenization strategies and vocabulary sizes on the performance of Arabic language models in downstream natural language processing tasks.
-
13 Feb 2024 1 repository listedMachine learning models for text classification often excel on in-distribution (ID) data but struggle with unseen out-of-distribution (OOD) inputs.
-
5 Feb 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedIn human-centric settings like education or healthcare, model accuracy and model explainability are key factors for user adoption.
-
30 Aug 2023 1 repository listedKyrgyz is a very underrepresented language in terms of modern natural language processing resources.
-
20 Jul 2023 1 repository listedPre-trained models for Czech Natural Language Processing are often evaluated on purely linguistic tasks (POS tagging, parsing, NER) and relatively simple classification tasks such as sentiment classification or article…
-
19 Apr 2023 1 repository listedFurthermore, we explore several alternatives to full fine-tuning of language models that are better suited for zero-shot and few-shot learning such as cross-lingual parameter-efficient fine-tuning (like MAD-X), pattern…
-
12 Dec 2022 1 repository listedWith the long-term goal of understanding how language is used and evolves within online communities, this work explores the application of natural language processing techniques to classify text articles according to…
-
25 Nov 2022 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedTemporal shifts -- distribution shifts arising from the passage of time -- often occur gradually and have the additional structure of timestamp metadata.
-
25 Nov 2022 1 repository listedIn this work, we propose Multiverse -- a new feature based on multilingual evidence that can be used for fake news detection and improve existing approaches.
-
14 Jun 2022 1 repository listedIn addition, we propose a self-supervised learning strategy based on SRLP to enhance the out-of-distribution generalization performance of our system.
-
29 Jan 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedWe empirically measure the effect of transliteration on MLLMs in this context.
-
20 Sep 2021 1 repository listedExperiments over the ISOT and Combined Corpus datasets show that transformers achieve an increase in F1 scores of up to 4.
-
30 Aug 2021 1 repository listedCurrent news datasets merely focus on text features on the news and rarely leverage the feature of images, excluding numerous essential features for news classification.
-
1 Aug 2021 1 repository listedMisleading information spreads on the Internet at an incredible speed, which can lead to irreparable consequences in some cases.
-
14 Jun 2021 1 repository listedHowever, given a large text corpus, representing all the words is not efficient in terms of vocabulary size.
-
7 May 2021 1 repository listedThis paper releases "AraCOVID19-MFH" a manually annotated multi-label Arabic COVID-19 fake news and hate speech detection dataset.
-
16 Mar 2021 1 repository listedThis work empirically demonstrates the ability of Text Graph Convolutional Network (Text GCN) to outperform traditional natural language processing benchmarks for the task of semi-supervised Swahili news classification.
-
8 Nov 2020 1 repository listedThese resources include: (a) large-scale sentence-level monolingual corpora, (b) pre-trained word embeddings, (c) pre-trained language models, and (d) multiple NLU evaluation datasets (IndicGLUE benchmark).
-
23 Oct 2020 1 repository listedRecent progress in text classification has been focused on high-resource languages such as English and Chinese.
-
1 Oct 2020 1 repository listedNeural network NLP models are vulnerable to small modifications of the input that maintain the original meaning but result in a different prediction.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections