Home › Datasets › task › Language Modelling
Language Modelling datasets
archive 2025-07-28
166 datasets carry the task tag "Language Modelling" (the task itself: Language Modelling), ordered by the archive's paper count. Page 3 of 4: 48 shown of 166. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Language Modelling datasets 97–144 of 166
UKP (UKP Argument Annotated Essays)
The UKP Argument Annotated Essays corpus consists of argument annotated persuasive essays including annotations of argument components and argumentative relations.
8 papers · 0 benchmarks
News translation is a recurring WMT task.
8 papers · 0 benchmarks
Winogender Schemas is a novel, Winograd schema-style set of minimal pair sentences that differ only by pronoun gender.
8 papers · 0 benchmarks
Composes sentence pairs (i.e., twin sentences).
7 papers · 0 benchmarks
GINC (Generative IN-Context learning Dataset)
GINC (Generative In-Context learning Dataset) is a small-scale synthetic dataset for studying in-context learning.
7 papers · 0 benchmarks
TUT-SED Synthetic 2016 contains of mixture signals artificially generated from isolated sound events samples.
7 papers · 0 benchmarks
CLUECorpus2020 is a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation.
6 papers · 0 benchmarks
The CoarseWSD-20 dataset is a coarse-grained sense disambiguation dataset built from Wikipedia (nouns only) targeting 2 to 5 senses of 20 ambiguous words.
6 papers · 0 benchmarks
A quantitative benchmark for developing and understanding video of fill-in-the-blank question-answering dataset with over 300,000 examples, based on descriptive video annotations for the visually impaired.
6 papers · 0 benchmarks
WNLaMPro (WordNet Language Model Probing)
The WordNet Language Model Probing (WNLaMPro) dataset consists of relations between keywords and words.
6 papers · 0 benchmarks
An unsupervised dataset for co-reference resolution.
6 papers · 0 benchmarks
Benchmark dataset for abstracts and titles of 100,000 ArXiv scientific papers.
6 papers · 1 benchmark
Coached Conversational Preference Elicitation is a dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
5 papers · 0 benchmarks
A large-scale Indonesian summarization dataset consisting of harvested articles from Liputan6.com, an online news portal, resulting in 215,827 document-summary pairs.
5 papers · 0 benchmarks
Tencent ML-Images is a large open-source multi-label image database, including 17,609,752 training and 88,739 validation image URLs, which are annotated with up to 11,166 categories.
5 papers · 0 benchmarks
The corpus represents the largest existing corpus of Catalan containing 687 million words, which is a significant increase given that until now the biggest corpus of Catalan, CuCWeb, counts 166 million words.
5 papers · 0 benchmarks
CCPE-M (Coached Conversational Preference Elicitation dataset for Movies)
A dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
4 papers · 0 benchmarks
CommitChronicle is a dataset for commit message generation (and/or completion).
4 papers · 0 benchmarks
Databricks Dolly 15k is a dataset containing 15,000 high-quality human-generated prompt / response pairs specifically designed for instruction tuning large language models.
4 papers · 0 benchmarks
Models character profiles and gives dialogue agents the ability to learn characters' language styles through their HLAs.
4 papers · 0 benchmarks
RONEC (Romanian Named Entity Corpus)
Romanian Named Entity Corpus is a named entity corpus for the Romanian language.
4 papers · 0 benchmarks
SLING consists of 38K minimal sentence pairs in Mandarin Chinese grouped into 9 high-level linguistic phenomena.
4 papers · 0 benchmarks
VietMed (VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain)
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
4 papers · 2 benchmarks
This is a dataset for disentangling conversations on IRC, which is the task of identifying separate conversations in a single stream of messages.
4 papers · 3 benchmarks
ChrEn (Cherokee-English Parallel Dataset)
Cherokee-English Parallel Dataset is a low-resource dataset of 14,151 pairs of sentences with around 313K English tokens and 206K Cherokee tokens.
3 papers · 0 benchmarks
Multilingual text collection extracted from the Internet Archive and Common Crawl archives.
3 papers · 0 benchmarks
The IndicNLP corpus is a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families.
3 papers · 0 benchmarks
KMIR (Knowledge Memorization, Identification, and Reasoning)
KMIR (Knowledge Memorization, Identification, and Reasoning) is a benchmark that covers 3 types of knowledge, including general knowledge, domain-specific knowledge, and commonsense, and provides 184,348 well-designed questions.
3 papers · 0 benchmarks
SLNET (SLNET: A Redistributable Corpus of 3rd-party Simulink Models)
SLNET is collection of third party Simulink models.
3 papers · 0 benchmarks
Wiki-Convert is a 900,000+ sentences dataset of precise number annotations from English Wikipedia.
3 papers · 0 benchmarks
WikiText-TL-39 is a benchmark language modeling dataset in Filipino that has 39 million tokens in the training set.
3 papers · 0 benchmarks
The Circa (meaning ‘approximately’) dataset aims to help machine learning systems to solve the problem of interpreting indirect answers to polar questions.
2 papers · 0 benchmarks
Comparative Question Completion is a dataset to evaluate what do large Language Models learn.
2 papers · 0 benchmarks
Czech restaurant information is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language.
2 papers · 1 benchmark
NQuAD (Nuclear Question Answering Dataset)
NQuAD is a Nuclear Question Answering Dataset, which contains 700+ nuclear Question Answer pairs developed and verified by expert nuclear researchers.
2 papers · 0 benchmarks
A collection of 385,705 scientific abstracts about Cognitive Control and their GPT-3 embeddings.
2 papers · 1 benchmark
RTC is a benchmark corpus of social media comments sampled over three years.
2 papers · 0 benchmarks
S-TEST is a benchmark for measuring the specificity of the language of pre-trained language models.
2 papers · 0 benchmarks
SART is a collection of three datasets for Similarity, Analogies and Relatedness for the Tatar language.
2 papers · 0 benchmarks
Contents (As on March 4, 2019) -------- The text corpus contains running text from various free licensed sources.
2 papers · 0 benchmarks
The Sentimental LIAR dataset is a modified and further extended version of the LIAR extension introduced by Kirilin et al.
2 papers · 0 benchmarks
Spades (Semantic PArsing of DEclarative Sentences)
Datasets Spades contains 93,319 questions derived from clueweb09 sentences.
2 papers · 0 benchmarks
The Stack Exchange dataset is a collection of data from various Stack Exchange sites, including Stack Overflow, Mathematics, Super User, and many others.
2 papers · 1 benchmark
ViText2SQL is a dataset for the Vietnamese Text-to-SQL semantic parsing task, consisting of about 10K question and SQL query pairs.
2 papers · 0 benchmarks
The Alexa Point of View dataset is point of view conversion dataset, a parallel corpus of messages spoken to a virtual assistant and the converted messages for delivery.
1 paper · 1 benchmark
The Books3 dataset emerged as part of a broader effort to train AI models for natural language understanding and generation.
1 paper · 1 benchmark
The Curation Corpus is a collection of 40,000 professionally-written summaries of news articles, with links to the articles themselves.
1 paper · 1 benchmark
DpgMedia2019 is a Dutch news dataset for partisanship detection.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.