Home › Datasets › task › Language Modelling

Language Modelling datasets

archive 2025-07-28

166 datasets carry the task tag "Language Modelling" (the task itself: Language Modelling), ordered by the archive's paper count. Page 2 of 4: 48 shown of 166. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Language Modelling datasets 49–96 of 166

DART is a large dataset for open-domain structured data record to text generation.
45 papers · 3 benchmarks
MLSUM (MultiLingual SUMmarization)
A large-scale MultiLingual SUMmarization dataset.
45 papers · 4 benchmarks
EmotionLines contains a total of 29245 labeled utterances from 2000 dialogues.
44 papers · 1 benchmark
OCNLI (Original Chinese Natural Language Inference)
OCNLI stands for Original Chinese Natural Language Inference.
44 papers · 0 benchmarks
RWC (Real World Computing Music Database)
The RWC (Real World Computing) Music Database is a copyright-cleared music database (DB) that is available to researchers as a common foundation for research.
43 papers · 0 benchmarks
PearRead is a dataset of scientific peer reviews.
42 papers · 0 benchmarks
SMM4H (Social Media Mining for Health Shared Task)
Social Media Mining for Health (SMM4H) Shared Task is a massive data source for biomedical and public health applications.
41 papers · 0 benchmarks
Tatoeba is a free collection of example sentences with translations geared towards foreign language learners.
38 papers · 2 benchmarks
AdvGLUE (Adversarial GLUE)
Adversarial GLUE (AdvGLUE) is a new multi-task benchmark to quantitatively and thoroughly explore and evaluate the vulnerabilities of modern large-scale language models under various types of adversarial attacks.
36 papers · 1 benchmark
Arxiv HEP-TH (high energy physics theory) citation graph is from the e-print arXiv and covers all the citations within a dataset of 27,770 papers with 352,807 edges.
35 papers · 5 benchmarks
2000 HUB5 English Evaluation Transcripts was developed by the Linguistic Data Consortium (LDC) and consists of transcripts of 40 English telephone conversations used in the 2000 HUB5 evaluation sponsored by NIST (National Institute of…
33 papers · 2 benchmarks
Worldtree is a corpus of explanation graphs, explanatory role ratings, and associated tablestore.
32 papers · 0 benchmarks
A new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families containing round 40 billion characters and aimed to accelerate the research of multilingual modeling.
30 papers · 3 benchmarks
MassiveText is a collection of large English-language text datasets from multiple sources: web pages, books, news articles, and code.
29 papers · 0 benchmarks
MuseData is an electronic library of orchestral and piano classical music from CCARH.
28 papers · 0 benchmarks
KPTimes is a large-scale dataset of news texts paired with editor-curated keyphrases.
27 papers · 3 benchmarks
The SentiCap dataset contains several thousand images with captions with positive and negative sentiments.
26 papers · 0 benchmarks
KELM is a large-scale synthetic corpus of Wikidata KG as natural text.
24 papers · 0 benchmarks
The Natural Stories dataset consists of English texts edited to contain many low-frequency syntactic constructions while still sounding fluent to native speakers.
24 papers · 0 benchmarks
Do-Not-Answer is a dataset to evaluate safeguards in large language models, and deploy safer open-source LLMs at a low cost.
22 papers · 0 benchmarks
The ECGs in this collection were obtained using a non-commercial, PTB prototype recorder with the following specifications: 16 input channels, (14 for ECGs, 1 for respiration, 1 for line voltage) Input voltage: ±16 mV, compensated offset…
22 papers · 4 benchmarks
Desc: About of Text8
22 papers · 1 benchmark
CLOTH (CLOze test by TeacHers)
The Cloze Test by Teachers (CLOTH) benchmark is a collection of nearly 100,000 4-way multiple-choice cloze-style questions from middle- and high school-level English language exams, where the answer fills a blank in a given text.
20 papers · 0 benchmarks
There are now many computer programs for automatically determining the sense of a word in context (Word Sense Disambiguation or WSD).
19 papers · 0 benchmarks
Taskmaster-1 is a dialog dataset consisting of 13,215 task-based dialogs in English, including 5,507 spoken and 7,708 written dialogs created with two distinct procedures.
19 papers · 0 benchmarks
CASIA-HWDB is a dataset for handwritten Chinese character recognition.
17 papers · 0 benchmarks
BeerAdvocate is a dataset that consists of beer reviews from beeradvocate.
15 papers · 1 benchmark
CMU DoG (CMU Document Grounded Conversations Dataset)
This is a document grounded dataset for text conversations.
15 papers · 0 benchmarks
Humicroedit is a humorous headline dataset.
15 papers · 0 benchmarks
BabyLM is a dataset for small scale language modeling, human language acquisition, low-resource NLP, and cognitive modeling.
14 papers · 0 benchmarks
The Collaborative Drawing game (CoDraw) dataset contains ~10K dialogs consisting of ~138K messages exchanged between human players in the CoDraw game.
14 papers · 0 benchmarks
The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages.
14 papers · 0 benchmarks
The IndoNLU benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems for Bahasa Indonesia.
14 papers · 2 benchmarks
OVAD benchmark (Open-Vocabulary Attribute Detection)
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner.
14 papers · 3 benchmarks
Open-Platypus is a family of fine-tuned and merged Large Language Models (LLMs) that achieves the strongest performance and currently stands at first place in HuggingFace's Open LLM Leaderboard.
14 papers · 0 benchmarks
ANTIQUE is a collection of 2,626 open-domain non-factoid questions from a diverse set of categories.
13 papers · 0 benchmarks
PersonalDialog is a large-scale multi-turn dialogue dataset containing various traits from a large number of speakers.
13 papers · 0 benchmarks
Chinese Few-shot Learning Evaluation Benchmark (FewCLUE) is a comprehensive small sample evaluation benchmark in Chinese.
12 papers · 5 benchmarks
The Hutter Prize Wikipedia dataset, also known as enwiki8, is a byte-level dataset consisting of the first 100 million bytes of a Wikipedia XML dump.
12 papers · 1 benchmark
Housekeep a benchmark to evaluate common sense reasoning in the home for embodied AI.
11 papers · 0 benchmarks
CC-Stories (or STORIES) is a dataset for common sense reasoning and language modeling.
10 papers · 0 benchmarks
ART Dataset (Abductive Reasoning in narrative Text)
ART consists of over 20k commonsense narrative contexts and 200k explanations.
9 papers · 0 benchmarks
ComQA is a large dataset of real user questions that exhibit different challenging aspects such as compositionality, temporal reasoning, and comparisons.
9 papers · 0 benchmarks
The IndoSum dataset is a benchmark dataset for Indonesian text summarization.
9 papers · 0 benchmarks
The SALMon dataset and benchmark was introduced in the paper "A Suite for Acoustic Language Model Evaluation", with the goal of evaluating the modelling abilities of speech language models with regards to different kinds of acoustic…
9 papers · 1 benchmark
Peyma is a Persian NER dataset to train and test NER systems.
8 papers · 0 benchmarks
The TUT Sound Events 2017 dataset contains 24 audio recordings in a street environment and contains 6 different classes.
8 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.