Home › Datasets › task › Natural Language Inference

Natural Language Inference datasets

archive 2025-07-28

81 datasets carry the task tag "Natural Language Inference" (the task itself: Natural Language Inference), ordered by the archive's paper count. Page 1 of 2: 48 shown of 81. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Natural Language Inference datasets 1–48 of 81

GLUE (General Language Understanding Evaluation benchmark)
General Language Understanding Evaluation (GLUE) benchmark is a collection of nine natural language understanding tasks, including single-sentence tasks CoLA and SST-2, similarity and paraphrasing tasks MRPC, STS-B and QQP, and natural…
3,197 papers · 13 benchmarks
MultiNLI (Multi-Genre Natural Language Inference)
The Multi-Genre Natural Language Inference (MultiNLI) dataset has 433K sentence pairs.
1,830 papers · 4 benchmarks
SNLI (Stanford Natural Language Inference)
The SNLI dataset (Stanford Natural Language Inference) consists of 570k sentence-pairs manually labeled as entailment, contradiction, and neutral.
1,311 papers · 1 benchmark
QNLI (Question-answering NLI)
The QNLI (Question-answering NLI) dataset is a Natural Language Inference dataset automatically derived from the Stanford Question Answering Dataset v1.1 (SQuAD).
1,234 papers · 3 benchmarks
MRPC (Microsoft Research Paraphrase Corpus)
Microsoft Research Paraphrase Corpus (MRPC) is a corpus consists of 5,801 sentence pairs collected from newswire articles.
786 papers · 4 benchmarks
WinoGrande is a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset.
703 papers · 7 benchmarks
XNLI (Cross-lingual Natural Language Inference)
The Cross-lingual Natural Language Inference (XNLI) corpus is the extension of the Multi-Genre NLI (MultiNLI) corpus to 15 languages.
349 papers · 7 benchmarks
SICK (Sentences Involving Compositional Knowledge)
The Sentences Involving Compositional Knowledge (SICK) dataset is a dataset for compositional distributional semantics.
348 papers · 5 benchmarks
ANLI (Adversarial NLI)
The Adversarial Natural Language Inference (ANLI, Nie et al.) is a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure.
287 papers · 3 benchmarks
PAWS (Paraphrase Adversaries from Word Scrambling)
Paraphrase Adversaries from Word Scrambling (PAWS) is a dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of…
159 papers · 0 benchmarks
e-SNLI is used for various goals, such as obtaining full sentence justifications of a model's decisions, improving universal sentence representations and transferring to out-of-domain NLI datasets.
139 papers · 1 benchmark
TabFact is a large-scale dataset which consists of 117,854 manually annotated statements with regard to 16,573 Wikipedia tables, their relations are classified as ENTAILED and REFUTED.
129 papers · 2 benchmarks
Visual Entailment (VE) consists of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks.
117 papers · 2 benchmarks
GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of (ambiguous pronoun, antecedent name), sampled from Wikipedia and released by Google AI Language for the evaluation of coreference resolution in practical…
104 papers · 0 benchmarks
The ListOps examples are comprised of summary operations on lists of single digit integers, written in prefix notation.
97 papers · 0 benchmarks
AQUA-RAT (Algebra Question Answering with Rationales)
Algebra Question Answering with Rationales (AQUA-RAT) is a dataset that contains algebraic word problems with rationales.
64 papers · 0 benchmarks
RTE (Recognizing Textual Entailment)
The Recognizing Textual Entailment (RTE) datasets come from a series of textual entailment challenges.
56 papers · 2 benchmarks
Quora Question Pairs (QQP) dataset consists of over 400,000 question pairs, and each question pair is annotated with a binary value indicating whether the two questions are paraphrase of each other.
55 papers · 8 benchmarks
OCNLI (Original Chinese Natural Language Inference)
OCNLI stands for Original Chinese Natural Language Inference.
44 papers · 0 benchmarks
SCROLLS (Standardized CompaRison Over Long Language Sequences)
SCROLLS (Standardized CompaRison Over Long Language Sequences) is an NLP benchmark consisting of a suite of tasks that require reasoning over long texts.
42 papers · 1 benchmark
Break is a question understanding dataset, aimed at training models to reason over complex questions.
39 papers · 0 benchmarks
MedQuAD (Medical Question Answering Dataset)
MedQuAD includes 47,457 medical question-answer pairs created from 12 NIH websites (e.g.
32 papers · 0 benchmarks
ContractNLI is a dataset for document-level natural language inference (NLI) on contracts whose goal is to automate/support a time-consuming procedure of contract review.
31 papers · 0 benchmarks
The HELP dataset is an automatically created natural language inference (NLI) dataset that embodies the combination of lexical and logical inferences focusing on monotonicity (i.e., phrase replacement-based reasoning).
30 papers · 1 benchmark
Chaos NLI is a Natural Language Inference (NLI) dataset with 100 annotations per example (for a total of 464,500 annotations) for some existing data points in the development sets of SNLI, MNLI, and Abductive NLI.
27 papers · 0 benchmarks
HeadQA is a multi-choice question answering testbed to encourage research on complex reasoning.
23 papers · 1 benchmark
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
KLUE (Korean Language Understanding Evaluation)
Korean Language Understanding Evaluation (KLUE) benchmark is a series of datasets to evaluate natural language understanding capability of Korean language models.
21 papers · 1 benchmark
QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms.
20 papers · 0 benchmarks
Torque is an English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships.
19 papers · 1 benchmark
InfoTabS comprises of human-written textual hypotheses based on premises that are tables extracted from Wikipedia info-boxes.
18 papers · 0 benchmarks
KorNLI is a Korean Natural Language Inference (NLI) dataset.
18 papers · 0 benchmarks
MED (Monotonicity Entailment Dataset)
MED is a new evaluation dataset that covers a wide range of monotonicity reasoning that was created by crowdsourcing and collected from linguistics publications.
18 papers · 1 benchmark
IndicGLUE (Indic General Language Understanding Evaluation Benchmark)
We now introduce IndicGLUE, the Indic General Language Understanding Evaluation Benchmark, which is a collection of various NLP tasks as de- scribed below.
16 papers · 4 benchmarks
WNLI (Winograd NLI)
The WNLI dataset is a part of the GLUE benchmark used for Natural Language Inference (NLI).
14 papers · 1 benchmark
KorSTS is a dataset for semantic textural similarity (STS) in Korean.
13 papers · 0 benchmarks
Story Commonsense is a new large-scale dataset with rich low-level annotations and establishes baseline performance on several new tasks, suggesting avenues for future research.
13 papers · 0 benchmarks
FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
TaxiNLI is a dataset collected based on the principles and categorizations of the aforementioned taxonomy.
11 papers · 0 benchmarks
The CommitmentBank is a corpus of 1,200 naturally occurring discourses whose final sentence contains a clause-embedding predicate under an entailment canceling operator (question, modal, negation, antecedent of conditional).
10 papers · 1 benchmark
NLI4CT dataset consists of 2,400 annotated statements with accompanying labels, CTRs, and evidence.
10 papers · 0 benchmarks
ART Dataset (Abductive Reasoning in narrative Text)
ART consists of over 20k commonsense narrative contexts and 200k explanations.
9 papers · 0 benchmarks
Natural Language Inference (NLI), also called Textual Entailment, is an important task in NLP with the goal of determining the inference relationship between a premise p and a hypothesis h.
9 papers · 1 benchmark
MedNLI (Medical Natural Language Inference)
The MedNLI dataset consists of the sentence pairs developed by Physicians from the Past Medical History section of MIMIC-III clinical notes annotated for Definitely True, Maybe True and Definitely False.
9 papers · 2 benchmarks
The SciTail dataset is an entailment dataset created from multiple-choice science exams and web sentences.
9 papers · 1 benchmark
SelQA is a dataset that consists of questions generated through crowdsourcing and sentence length answers that are drawn from the ten most prevalent topics in the English Wikipedia.
9 papers · 0 benchmarks
DocNLI is a large-scale dataset for document-level NLI.
8 papers · 0 benchmarks
DaNetQA (Yes/no Question Answering Dataset for the Russian)
DaNetQA is a question answering dataset for yes/no questions.
7 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.