Datasets › Yahoo! Answers

Yahoo! Answers

Introduced by Xiang Zhang et al. in Character-level Convolutional Networks for Text Classification4 Sep 2015 archive 2025-07-28

The Yahoo! Answers topic classification dataset is constructed using 10 largest main categories. Each class contains 140,000 training samples and 6,000 testing samples. Therefore, the total number of training samples is 1,400,000 and testing samples 60,000 in this dataset. From all the answers and other meta-information, we only used the best answer content and the main category information. Source:github

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text Classification Yahoo! Answers BERT-ITPT-FiT Accuracy 77.62 How to Fine-Tune BERT for Text Classification? xuyige/BERT4doc-Classification +14 10 Compare
Unsupervised Text Classification Yahoo! Answers Lbl2TransformerVec F1-score 55.84 Evaluating Unsupervised Text Classification: Zero-shot... sebischair/lbl2vec +1 1 Compare

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 132. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Evaluating Unsupervised Text Classification: Zero-shot and Similarity-based Approaches 2 1 29 Nov 2022 not harvested
Sampling Bias in Deep Active Classification: An Empirical Study 2 1 20 Sep 2019 ran 0 of 1 samples (1 unverified)
DELTA: A DEep learning based Language Technology plAtform 2 1 2 Aug 2019 not harvested
How to Fine-Tune BERT for Text Classification? 15 1 14 May 2019 ran 6 of 18 samples (12 unverified; 5 pointer-only for licence)
Learning to Remember More with Less Memorization 1 1 5 Jan 2019 ran 3 of 3 samples (0 unverified)
Explicit Interaction Model towards Text Classification 1 1 23 Nov 2018 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
Compositional Coding Capsule Network with K-Means Routing for Text Classification 1 1 22 Oct 2018 not harvested
Disconnected Recurrent Neural Networks for Text Categorization 0 1 1 Jul 2018 not harvested
Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling Mechanisms 2 1 24 May 2018 not harvested
Abstractive Text Classification Using Sequence-to-convolution Neural Networks 1 1 20 May 2018 not harvested
Bag of Tricks for Efficient Text Classification 65 1 6 Jul 2016 ran 2 of 9 samples (7 unverified; 2 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Yahoo! Answers

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections