Home › Datasets › task › Natural Language Inference
Natural Language Inference datasets
archive 2025-07-28
81 datasets carry the task tag "Natural Language Inference" (the task itself: Natural Language Inference), ordered by the archive's paper count. Page 2 of 2: 33 shown of 81. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Natural Language Inference datasets 49–81 of 81
JGLUE, Japanese General Language Understanding Evaluation, is built to measure the general NLU ability in Japanese.
7 papers · 0 benchmarks
KUAKE Query-Query Relevance, a dataset used to evaluate the relevance of the content expressed in two queries, is used for the KUAKE-QQR task.
7 papers · 1 benchmark
KUAKE Query Title Relevance, a dataset used to estimate the relevance of the title of a query document, is used for the KUAKE-QTR task.
7 papers · 1 benchmark
TERRa (Textual Entailment Recognition for Russian)
Textual Entailment Recognition has been proposed recently as a generic task that captures major semantic inference needs across many NLP applications, such as Question Answering, Information Retrieval, Information Extraction, and Text…
7 papers · 1 benchmark
An IMPlicature and PRESupposition diagnostic dataset (IMPPRES), consisting of >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types.
6 papers · 0 benchmarks
RCB (Russian Commitment Bank)
The Russian Commitment Bank is a corpus of naturally occurring discourses whose final sentence contains a clause-embedding predicate under an entailment cancelling operator (question, modal, negation, antecedent of conditional).
6 papers · 1 benchmark
BioNLI (Biomedical Natural Language Inference)
BioNLI is a dataset in biomedical natural language inference.
5 papers · 1 benchmark
LiDiRus (Linguistic Diagnostic for Russian)
LiDiRus is a diagnostic dataset that covers a large volume of linguistic phenomena, while allowing you to evaluate information systems on a simple test of textual entailment recognition.
5 papers · 1 benchmark
RuSentRel is a corpus of analytical articles translated into Russian texts in the domain of international politics obtained from foreign authoritative sources.
5 papers · 0 benchmarks
IndoNLI is the first human-elicited NLI dataset for Indonesian consisting of nearly 18K sentence pairs annotated by crowd workers and experts.
4 papers · 0 benchmarks
We generate epistemic reasoning problems using modal logic to target theory of mind (tom) in natural language processing models.
4 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
JUSTICE (JUSTICE: A Dataset for Supreme Court’s Judgment Prediction)
The dataset contains 3304 cases from the Supreme Court of the United States from 1955 to 2021.
3 papers · 0 benchmarks
JaNLI (Japanese Adversarial Natural Language Inference)
The Japanese Adversarial NLI (JaNLI) dataset is designed to require understanding of Japanese linguistic phenomena and illuminate the vulnerabilities of models.
3 papers · 0 benchmarks
Pars-ABSA is a manually annotated Persian dataset, Pars-ABSA, which is verified by 3 native Persian speakers.
3 papers · 0 benchmarks
Huggingface Datasets is a great library, but it lacks standardization, and datasets require preprocessing work to be used interchangeably.
3 papers · 0 benchmarks
WiLI-2018 is a benchmark dataset for monolingual written natural language identification.
3 papers · 0 benchmarks
esXNLI is a bilingual NLI dataset.
3 papers · 0 benchmarks
Natural Language Inference processes pairs of sentences to extract their semantic relations.
2 papers · 0 benchmarks
NLI-TR (Natural Language Inference in Turkish)
Natural Language Inference in Turkish (NLI-TR) provides translations of two large English NLI datasets into Turkish and had a team of experts validate their translation quality and fidelity to the original labels.
2 papers · 0 benchmarks
NewsPH-NLI is a sentence entailment benchmark dataset in the low-resource Filipino language.
2 papers · 0 benchmarks
PropSegmEnt is a corpus of over 35K propositions annotated by expert human raters.
2 papers · 0 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
This dataset is named as the DistNLI dataset, which is a synthesized benchmark aiming to probe neural network models from the aspect of conjunctions on distributivity in NLI task in American English.
1 paper · 0 benchmarks
GD-NLI (Generated Debiased NLI Datasets)
This is a set of debiased Natural Language Inference (NLI) datasets produced by the paper Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets.
1 paper · 0 benchmarks
GQNLI-FR is a manually translated French version of the GQNLI challenge dataset, originally written in English.
1 paper · 0 benchmarks
The Gigaword Entailment dataset is a dataset for entailment prediction between an article and its headline.
1 paper · 0 benchmarks
HANS (Heuristic Analysis for NLI Systems)
The HANS (Heuristic Analysis for NLI Systems) dataset which contains many examples where the heuristics fail.
1 paper · 1 benchmark
JamPatoisNLI provides the first dataset for natural language inference in a creole language, Jamaican Patois.
1 paper · 1 benchmark
NLI4Wills Corpus can be used to train transformers and sentence-transformer models for the validity evaluation of the legal will statements.
1 paper · 0 benchmarks
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP), e.g.
1 paper · 1 benchmark
RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).
1 paper · 0 benchmarks
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP, also called verbal probabilities), e.g.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.