Home › Datasets › task › Natural Language Inference

Natural Language Inference datasets

archive 2025-07-28

81 datasets carry the task tag "Natural Language Inference" (the task itself: Natural Language Inference), ordered by the archive's paper count. Page 2 of 2: 33 shown of 81. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Natural Language Inference datasets 49–81 of 81

JGLUE, Japanese General Language Understanding Evaluation, is built to measure the general NLU ability in Japanese.
7 papers · 0 benchmarks
KUAKE-QQR (Query-Query Relevance Dataset)
KUAKE Query-Query Relevance, a dataset used to evaluate the relevance of the content expressed in two queries, is used for the KUAKE-QQR task.
7 papers · 1 benchmark
KUAKE-QTR (Query-Title Relevance Dataset)
KUAKE Query Title Relevance, a dataset used to estimate the relevance of the title of a query document, is used for the KUAKE-QTR task.
7 papers · 1 benchmark
TERRa (Textual Entailment Recognition for Russian)
Textual Entailment Recognition has been proposed recently as a generic task that captures major semantic inference needs across many NLP applications, such as Question Answering, Information Retrieval, Information Extraction, and Text…
7 papers · 1 benchmark
An IMPlicature and PRESupposition diagnostic dataset (IMPPRES), consisting of >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types.
6 papers · 0 benchmarks
RCB (Russian Commitment Bank)
The Russian Commitment Bank is a corpus of naturally occurring discourses whose final sentence contains a clause-embedding predicate under an entailment cancelling operator (question, modal, negation, antecedent of conditional).
6 papers · 1 benchmark
BioNLI (Biomedical Natural Language Inference)
BioNLI is a dataset in biomedical natural language inference.
5 papers · 1 benchmark
LiDiRus (Linguistic Diagnostic for Russian)
LiDiRus is a diagnostic dataset that covers a large volume of linguistic phenomena, while allowing you to evaluate information systems on a simple test of textual entailment recognition.
5 papers · 1 benchmark
RuSentRel is a corpus of analytical articles translated into Russian texts in the domain of international politics obtained from foreign authoritative sources.
5 papers · 0 benchmarks
IndoNLI is the first human-elicited NLI dataset for Indonesian consisting of nearly 18K sentence pairs annotated by crowd workers and experts.
4 papers · 0 benchmarks
We generate epistemic reasoning problems using modal logic to target theory of mind (tom) in natural language processing models.
4 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
JUSTICE (JUSTICE: A Dataset for Supreme Court’s Judgment Prediction)
The dataset contains 3304 cases from the Supreme Court of the United States from 1955 to 2021.
3 papers · 0 benchmarks
JaNLI (Japanese Adversarial Natural Language Inference)
The Japanese Adversarial NLI (JaNLI) dataset is designed to require understanding of Japanese linguistic phenomena and illuminate the vulnerabilities of models.
3 papers · 0 benchmarks
Pars-ABSA is a manually annotated Persian dataset, Pars-ABSA, which is verified by 3 native Persian speakers.
3 papers · 0 benchmarks
Huggingface Datasets is a great library, but it lacks standardization, and datasets require preprocessing work to be used interchangeably.
3 papers · 0 benchmarks
WiLI-2018 is a benchmark dataset for monolingual written natural language identification.
3 papers · 0 benchmarks
esXNLI is a bilingual NLI dataset.
3 papers · 0 benchmarks
Natural Language Inference processes pairs of sentences to extract their semantic relations.
2 papers · 0 benchmarks
NLI-TR (Natural Language Inference in Turkish)
Natural Language Inference in Turkish (NLI-TR) provides translations of two large English NLI datasets into Turkish and had a team of experts validate their translation quality and fidelity to the original labels.
2 papers · 0 benchmarks
NewsPH-NLI is a sentence entailment benchmark dataset in the low-resource Filipino language.
2 papers · 0 benchmarks
PropSegmEnt is a corpus of over 35K propositions annotated by expert human raters.
2 papers · 0 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
This dataset is named as the DistNLI dataset, which is a synthesized benchmark aiming to probe neural network models from the aspect of conjunctions on distributivity in NLI task in American English.
1 paper · 0 benchmarks
GD-NLI (Generated Debiased NLI Datasets)
This is a set of debiased Natural Language Inference (NLI) datasets produced by the paper Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets.
1 paper · 0 benchmarks
GQNLI-FR is a manually translated French version of the GQNLI challenge dataset, originally written in English.
1 paper · 0 benchmarks
The Gigaword Entailment dataset is a dataset for entailment prediction between an article and its headline.
1 paper · 0 benchmarks
HANS (Heuristic Analysis for NLI Systems)
The HANS (Heuristic Analysis for NLI Systems) dataset which contains many examples where the heuristics fail.
1 paper · 1 benchmark
JamPatoisNLI provides the first dataset for natural language inference in a creole language, Jamaican Patois.
1 paper · 1 benchmark
NLI4Wills Corpus can be used to train transformers and sentence-transformer models for the validity evaluation of the legal will statements.
1 paper · 0 benchmarks
Probability words NLI (Natural language inference with words estimative of probability (WEP))
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP), e.g.
1 paper · 1 benchmark
RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).
1 paper · 0 benchmarks
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP, also called verbal probabilities), e.g.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.