Datasets › ANLI

ANLI (Adversarial NLI)

Introduced by Yixin Nie et al. in Adversarial NLI: A New Benchmark for Natural Language Understanding1 Jan 2019 archive 2025-07-28

The Adversarial Natural Language Inference (ANLI, Nie et al.) is a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure. Particular, the data is selected to be difficult to the state-of-the-art models, including BERT and RoBERTa.

Source: The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding Image Source: https://arxiv.org/pdf/1910.14599.pdf

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Natural Language Inference ANLI test T5-3B (explanation prompting) A1 81.8 Prompting for explanations improves Adversarial NLI. Is... — 25 Compare
Natural Language Inference ANLI no rows — — 0 Compare
Natural Language Inference ANLI-r3 no rows — — 0 Compare

Papers archive 2025-07-28

14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 287. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets 1 1 29 May 2023 not harvested
The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning 2 1 23 May 2023 not harvested
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
Prompting for explanations improves Adversarial NLI. Is this true? {Yes} it is {true} because {it weakens superficial cues} 0 2 1 May 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
Exploring the Benefits of Training Expert Language Models over Instruction Tuning 2 1 7 Feb 2023 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models 0 1 28 Oct 2022 not harvested
Large Language Models Can Self-Improve 0 6 20 Oct 2022 not harvested
Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners 1 1 6 Oct 2022 not harvested
InfoBERT: Improving Robustness of Language Models from An Information Theoretic Perspective 2 1 5 Oct 2020 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
Language Models are Few-Shot Learners 67 1 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Adversarial Training for Large Neural Language Models 3 1 20 Apr 2020 not harvested
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
XLNet: Generalized Autoregressive Pretraining for Language Understanding 27 1 19 Jun 2019 ran 10 of 24 samples (14 unverified; 3 pointer-only for licence)

Dataset loaders archive 2025-07-28

7 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ANLI test
  • ANLI
  • ANLI-all
  • ANLI-r3

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections