Home › Datasets › task › Language Modelling

Language Modelling datasets

archive 2025-07-28

166 datasets carry the task tag "Language Modelling" (the task itself: Language Modelling), ordered by the archive's paper count. Page 1 of 4: 48 shown of 166. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Language Modelling datasets 1–48 of 166

The IMDb Movie Reviews dataset is a binary sentiment analysis dataset consisting of 50,000 reviews from the Internet Movie Database (IMDb) labeled as positive or negative.
1,787 papers · 9 benchmarks
The PubMed dataset consists of 19717 scientific publications from PubMed database pertaining to diabetes classified into one of three classes.
1,236 papers · 19 benchmarks
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia.
1,081 papers · 1 benchmark
The English Penn Treebank (PTB) corpus, and in particular the section of the corpus corresponding to the articles of Wall Street Journal (WSJ), is one of the most known and used corpus for the evaluation of models for sequence labelling.
1,006 papers · 10 benchmarks
C4 (Colossal Clean Crawled Corpus)
C4 is a colossal, cleaned version of Common Crawl's web crawl corpus.
981 papers · 1 benchmark
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia.
560 papers · 2 benchmarks
The Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together.
467 papers · 1 benchmark
BIG-bench (Beyond the Imitation Game Benchmark)
The Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark intended to probe large language models and extrapolate their future capabilities.
349 papers · 121 benchmarks
BookCorpus is a large collection of free novel books written by unpublished authors, which contains 11,038 books (around 74M sentences and 1G words) of 16 different sub-genres (e.g., Romance, Historical, Adventure, etc.).
344 papers · 1 benchmark
The LAMBADA (LAnguage Modeling Broadened to Account for Discourse Aspects) benchmark is an open-ended cloze task which consists of about 10,000 passages from BooksCorpus where a missing target word is predicted in the last sentence of each…
293 papers · 1 benchmark
The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts.
276 papers · 3 benchmarks
LAMA (LAnguage Model Analysis)
LAnguage Model Analysis (LAMA) consists of a set of knowledge sources, each comprised of a set of facts.
214 papers · 0 benchmarks
OpenSubtitles is collection of multilingual parallel corpora.
214 papers · 3 benchmarks
OpenWebText is an open-source recreation of the WebText corpus.
207 papers · 2 benchmarks
WiC (Words in Context)
WiC is a benchmark for the evaluation of context-sensitive word embeddings.
206 papers · 3 benchmarks
Clotho is an audio captioning dataset, consisting of 4981 audio samples, and each audio sample has five captions (a total of 24 905 captions).
202 papers · 3 benchmarks
AISHELL-1 is a corpus for speech recognition research and building speech recognition systems for Mandarin.
197 papers · 1 benchmark
Libri-Light is a collection of spoken English audio suitable for training speech recognition systems under limited or no supervision.
194 papers · 2 benchmarks
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance.
190 papers · 1 benchmark
Dataset of hate speech annotated on Internet forum posts in English at sentence-level.
180 papers · 1 benchmark
LRA (Long-Range Arena)
Long-range arena (LRA) is an effort toward systematic evaluation of efficient transformer models.
180 papers · 1 benchmark
PAWS-X contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean.
172 papers · 0 benchmarks
ATOMIC is an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge.
170 papers · 0 benchmarks
ELI5 is a dataset for long-form question answering.
158 papers · 1 benchmark
OLID (Offensive Language Identification Dataset)
The OLID is a hierarchical dataset to identify the type and the target of offensive texts in social media.
152 papers · 1 benchmark
A new open-vocabulary language modelling benchmark derived from books.
147 papers · 1 benchmark
The One Billion Word dataset is a dataset for language modeling.
141 papers · 2 benchmarks
BLUE (Biomedical Language Understanding Evaluation)
The BLUE benchmark consists of five different biomedicine text-mining tasks with ten corpora.
133 papers · 0 benchmarks
BLiMP (Benchmark of Linguistic Minimal Pairs)
BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English.
126 papers · 0 benchmarks
SIQA (Social Interaction QA)
Social Interaction QA (SIQA) is a question-answering benchmark for testing social commonsense intelligence.
120 papers · 1 benchmark
WritingPrompts is a large dataset of 300K human-written stories paired with writing prompts from an online forum.
118 papers · 1 benchmark
A dataset of large scale alignments between Wikipedia abstracts and Wikidata triples.
117 papers · 1 benchmark
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
CLUE (Chinese Language Understanding Evaluation Benchmark)
CLUE is a Chinese Language Understanding Evaluation benchmark.
99 papers · 8 benchmarks
E2E (End-to-End NLG Challenge)
End-to-End NLG Challenge (E2E) aims to assess whether recent end-to-end NLG systems can generate more complex output by learning from datasets containing higher lexical richness, syntactic complexity and diverse discourse phenomena.
90 papers · 4 benchmarks
KP20k is a large-scale scholarly articles dataset with 528K articles for training, 20K articles for validation and 20K articles for testing.
87 papers · 3 benchmarks
TweetEval introduces an evaluation framework consisting of seven heterogeneous Twitter-specific classification tasks.
84 papers · 1 benchmark
MetaQA (MoviE Text Audio QA)
The MetaQA dataset consists of a movie ontology derived from the WikiMovies Dataset and three sets of question-answer pairs written in natural language: 1-hop, 2-hop, and 3-hop queries.
81 papers · 1 benchmark
RealNews is a large corpus of news articles from Common Crawl.
80 papers · 0 benchmarks
The Semantic Scholar corpus (S2) is composed of titles from scientific papers published in machine learning conferences and journals from 1985 to 2017, split by year (33 timesteps).
75 papers · 0 benchmarks
CMRC (Chinese Machine Reading Comprehension)
CMRC is a dataset is annotated by human experts with near 20,000 questions as well as a challenging set which is composed of the questions that need reasoning over multiple clues.
69 papers · 0 benchmarks
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
The TED-LIUM corpus consists of English-language TED talks.
64 papers · 2 benchmarks
SciDocs evaluation framework consists of a suite of evaluation tasks designed for document-level tasks.
57 papers · 2 benchmarks
Jericho is a learning environment for man-made Interactive Fiction (IF) games.
54 papers · 0 benchmarks
emrQA has 1 million question-logical form and 400,000+ questionanswer evidence pairs.
50 papers · 0 benchmarks
CMRC 2018 (Chinese Machine Reading Comprehension 2018)
CMRC 2018 is a dataset for Chinese Machine Reading Comprehension.
49 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.