Home › Datasets › task › Natural Language Understanding
Natural Language Understanding datasets
archive 2025-07-28
72 datasets carry the task tag "Natural Language Understanding" (the task itself: Natural Language Understanding), ordered by the archive's paper count. Page 2 of 2: 24 shown of 72. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Natural Language Understanding datasets 49–72 of 72
GeoGLUE (GeoGraphic Language Understanding Evaluation Benchmark)
GeoGLUE is a GeoGraphic Language Understanding Evaluation benchmark, which consists of six geographic text-related tasks, including geographic textual similarity on recall, geotagged geographic elements tagging, geographic composition…
3 papers · 0 benchmarks
JUSTICE (JUSTICE: A Dataset for Supreme Court’s Judgment Prediction)
The dataset contains 3304 cases from the Supreme Court of the United States from 1955 to 2021.
3 papers · 0 benchmarks
MuLD (Multitask Long Document Benchmark)
MuLD (Multitask Long Document Benchmark) is a set of 6 NLP tasks where the inputs consist of at least 10,000 words.
3 papers · 5 benchmarks
NusaCrowd is a collaborative initiative to collect and unite existing resources for Indonesian languages, including opening access to previously non-public resources.
3 papers · 0 benchmarks
WikiText-TL-39 is a benchmark language modeling dataset in Filipino that has 39 million tokens in the training set.
3 papers · 0 benchmarks
RuMedBench is a benchmark dataset for Russian medical language understanding.
2 papers · 0 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
CLUES (Constrained Language Understanding Evaluation Standard)
CLUES (Constrained Language Understanding Evaluation Standard) is a benchmark for evaluating the few-shot learning capabilities of NLU models.
1 paper · 0 benchmarks
ChatLog is a coarse-to-fine temporal dataset called ChatLog, consisting of two parts that update monthly and daily: 1.
1 paper · 0 benchmarks
Dialog-based Language Learning dataset is designed to measure how well models can perform at learning as a student given a teacher’s textual responses to the student’s answer (as well as potentially receiving an external real-valued reward…
1 paper · 0 benchmarks
Dataset Summary The DiscoEval is an English-language Benchmark that contains a test suite of 7 tasks to evaluate whether sentence representations include semantic information relevant to discourse processing.
1 paper · 0 benchmarks
EPIC30M contains a subset of 26.2 millions tweets related to three general diseases, namely Ebola, Cholera and Swine Flu, and another subset of 4.7 millions tweets of six global epidemic outbreaks, including 2009 H1N1 Swine Flu, 2010 Haiti…
1 paper · 0 benchmarks
ExPUNations is a humor dataset with such extensive and fine-grained annotations specifically for puns.
1 paper · 0 benchmarks
Introduction The FewGLUE64labeled dataset is a new version of FewGLUE dataset.
1 paper · 0 benchmarks
IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.
1 paper · 0 benchmarks
ImagiFilter focusses on photographic and/or natural images, a very common use-case in computer vision research.
1 paper · 0 benchmarks
MAVEN-Arg is an advanced event argument extraction dataset, which offers three main advantages: A comprehensive schema covering 162 event types and 612 argument roles, all with expert-written definitions and examples.
1 paper · 0 benchmarks
MultiWOZ-coref, (or MultiWOZ 2.3) is an extension of the MultiWOZ dataset that adds co-reference annotations in addition to corrections of dialogue acts and dialogue states.
1 paper · 0 benchmarks
This project is a collection of three corpora which can be used for evaluating chatbots or other conversational interfaces.
1 paper · 0 benchmarks
Phrase in Context is a curated benchmark for phrase understanding and semantic search, consisting of three tasks of increasing difficulty: Phrase Similarity (PS), Phrase Retrieval (PR) and Phrase Sense Disambiguation (PSD).
1 paper · 0 benchmarks
We release various types of word embeddings for multiple Indian languages.
1 paper · 0 benchmarks
Thunder-NUBench (Negation Understanding Benchmark) is a benchmark specifically designed to evaluate large language models’ (LLMs) sentence-level understanding of negation.
1 paper · 0 benchmarks
bigscience/P3 (bigscience/P3, split='ai2_arc_ARC_Challenge_pick_the_most_correct_option')
This datasets consists of challenging reasoning questions in multiple choice format.
1 paper · 0 benchmarks
IndirectRequests is an LLM-generated dataset of user utterances in a task-oriented dialogue setting where the user does not directly specify their preferred slot value.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.