Home › Datasets › task › Natural Language Understanding
Natural Language Understanding datasets
archive 2025-07-28
72 datasets carry the task tag "Natural Language Understanding" (the task itself: Natural Language Understanding), ordered by the archive's paper count. Page 1 of 2: 48 shown of 72. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Natural Language Understanding datasets 1–48 of 72
GLUE (General Language Understanding Evaluation benchmark)
General Language Understanding Evaluation (GLUE) benchmark is a collection of nine natural language understanding tasks, including single-sentence tasks CoLA and SST-2, similarity and paraphrasing tasks MRPC, STS-B and QQP, and natural…
3,197 papers · 13 benchmarks
MML (Massive Multitask Language Understanding)
MMLU (Massive Multitask Language Understanding) is a new benchmark designed to measure knowledge acquired during pretraining by evaluating models exclusively in zero-shot and few-shot settings.
1,922 papers · 29 benchmarks
WinoGrande is a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset.
703 papers · 7 benchmarks
SGD (Schema-Guided Dialogue)
The Schema-Guided Dialogue (SGD) dataset consists of over 20k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
186 papers · 2 benchmarks
LogiQA consists of 8,678 QA instances, covering multiple types of deductive reasoning.
127 papers · 1 benchmark
WritingPrompts is a large dataset of 300K human-written stories paired with writing prompts from an online forum.
118 papers · 1 benchmark
CosmosQA is a large-scale dataset of 35.6K problems that require commonsense-based reading comprehension, formulated as multiple-choice questions.
102 papers · 0 benchmarks
CLUE (Chinese Language Understanding Evaluation Benchmark)
CLUE is a Chinese Language Understanding Evaluation benchmark.
99 papers · 8 benchmarks
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants.
68 papers · 0 benchmarks
AQUA-RAT (Algebra Question Answering with Rationales)
Algebra Question Answering with Rationales (AQUA-RAT) is a dataset that contains algebraic word problems with rationales.
64 papers · 0 benchmarks
Legal General Language Understanding Evaluation (LexGLUE) benchmark is a collection of datasets for evaluating model performance across a diverse set of legal NLU tasks in a standardized way.
46 papers · 1 benchmark
A new publicly available dataset for verification of climate change-related claims.
38 papers · 1 benchmark
EmoBank is a corpus of 10k English sentences balancing multiple genres, annotated with dimensional emotion metadata in the Valence-Arousal-Dominance (VAD) representation format.
32 papers · 0 benchmarks
WikiCoref is an English corpus annotated for anaphoric relations, where all documents are from the English version of Wikipedia.
28 papers · 1 benchmark
decaNLP (Natural Language Decathlon Benchmark)
Natural Language Decathlon Benchmark (decaNLP) is a challenge that spans ten tasks: question answering, machine translation, summarization, natural language inference, sentiment analysis, semantic role labeling, zero-shot relation…
28 papers · 0 benchmarks
CrossWOZ is the first large-scale Chinese Cross-Domain Wizard-of-Oz task-oriented dataset.
25 papers · 0 benchmarks
MCScript is used as the official dataset of SemEval2018 Task11.
24 papers · 0 benchmarks
Moral Stories is a crowd-sourced dataset of structured narratives that describe normative and norm-divergent actions taken by individuals to accomplish certain intentions in concrete situations, and their respective consequences.
24 papers · 0 benchmarks
RecipeQA is a dataset for multimodal comprehension of cooking recipes.
24 papers · 1 benchmark
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
STREUSLE stands for Supersense-Tagged Repository of English with a Unified Semantics for Lexical Expressions.
21 papers · 1 benchmark
A large-scale English dataset for coreference resolution.
20 papers · 1 benchmark
KorNLI is a Korean Natural Language Inference (NLI) dataset.
18 papers · 0 benchmarks
DialoGLUE is a natural language understanding benchmark for task-oriented dialogue designed to encourage dialogue research in representation-based transfer, domain adaptation, and sample-efficient task learning.
17 papers · 2 benchmarks
The IndoNLU benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems for Bahasa Indonesia.
14 papers · 2 benchmarks
FewGLUE consists of a random selection of 32 training examples from the SuperGLUE training sets and up to 20,000 unlabeled examples for each SuperGLUE task.
13 papers · 0 benchmarks
JEC-QA is a LQA (Legal Question Answering) dataset collected from the National Judicial Examination of China.
13 papers · 0 benchmarks
The KLEJ benchmark (Kompleksowa Lista Ewaluacji Językowych) is a set of nine evaluation tasks for the Polish language understanding task.
13 papers · 0 benchmarks
KorSTS is a dataset for semantic textural similarity (STS) in Korean.
13 papers · 0 benchmarks
SARA (StAtutory Reasoning Assessment)
A dataset for statutory reasoning in tax law entailment and question answering.
13 papers · 0 benchmarks
Corpus containing 25206 sentences labelled with lexical instances of 717 idiomatic expressions.
12 papers · 0 benchmarks
Perspectrum is a dataset of claims, perspectives and evidence, making use of online debate websites to create the initial data collection, and augmenting it using search engines in order to expand and diversify the dataset.
11 papers · 1 benchmark
OCW (Only Connect Wall Dataset and creative problem solving tasks)
The OCW dataset is for evaluating creative problem solving tasks by curating the problems and human performance results from the popular British quiz show Only Connect.
10 papers · 1 benchmark
MMCU (Measuring Massive Multitask Chinese Understanding)
We propose a test to measure the multitask accuracy of large Chinese language models.
9 papers · 0 benchmarks
Source: BARThez: a Skilled Pretrained French Sequence-to-Sequence Model OrangeSum is a single-document extreme summarization dataset with two tasks: title and abstract.
8 papers · 1 benchmark
Composes sentence pairs (i.e., twin sentences).
7 papers · 0 benchmarks
The Medical Dataset for Abbreviation Disambiguation for Natural Language Understanding (MeDAL) is a large medical text dataset curated for abbreviation disambiguation, designed for natural language understanding pre-training in the medical…
7 papers · 0 benchmarks
An IMPlicature and PRESupposition diagnostic dataset (IMPPRES), consisting of >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types.
6 papers · 0 benchmarks
NLU++ (NLLU++ : A Multi-Label, Slot-Rich, Generalisable Dataset for Natural Language Understanding in Task-Oriented Dialogue)
nlu++ is a dataset for natural language understanding (NLU) in task-oriented dialogue (ToD) systems, with the aim to provide a much more challenging evaluation environment for dialogue NLU models, up to date with the current application…
6 papers · 0 benchmarks
VisPro dataset contains coreference annotation of 29,722 pronouns from 5,000 dialogues.
6 papers · 0 benchmarks
An unsupervised dataset for co-reference resolution.
6 papers · 0 benchmarks
A large-scale evaluation set that provides human ratings for the plausibility of 10,000 SP pairs over five SP relations, covering 2,500 most frequent verbs, nouns, and adjectives in American English.
5 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
A comprehensive multi-task benchmark for the Polish language understanding, accompanied by an online leaderboard.
3 papers · 0 benchmarks
A benchmark Arabic dataset for commonsense understanding and validation as well as a baseline research and models trained using the same dataset.
3 papers · 0 benchmarks
CHQ-Summ (Consumer Healthcare Question Summarization)
Contains 1507 domain-expert annotated consumer health questions and corresponding summaries.
3 papers · 0 benchmarks
A massive, deduplicated corpus of 7.4M Python files from GitHub.
3 papers · 0 benchmarks
Emotional Dialogue Acts data contains dialogue act labels for existing emotion multi-modal conversational datasets.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.