Datasets › MultiNLI

MultiNLI (Multi-Genre Natural Language Inference)

Introduced by Adina Williams et al. in A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference1 Jan 2018 archive 2025-07-28

The Multi-Genre Natural Language Inference (MultiNLI) dataset has 433K sentence pairs. Its size and mode of collection are modeled closely like SNLI. MultiNLI offers ten distinct genres (Face-to-face, Telephone, 9/11, Travel, Letters, Oxford University Press, Slate, Verbatim, Goverment and Fiction) of written and spoken English data. There are matched dev/test sets which are derived from the same sources as those in the training set, and mismatched sets which do not closely resemble any seen at training time.

Source: Semantic Sentence Matching with Densely-connectedRecurrent and Co-attentive Information

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Natural Language Inference MultiNLI Turing NLR v5 XXL 5.4B (fine-tuned) Matched 92.6 — — 67 Compare
Natural Language Inference MultiNLI Dev TinyBERT-6 67M Matched 84.5 TinyBERT: Distilling BERT for Natural Language Understanding PaddlePaddle/PaddleNLP +9 10 Compare
Natural Language Inference multi_nli no rows — — 0 Compare
Text Generation MNLI no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 40 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,830. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI 0 2 12 Dec 2024 not harvested
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale 2 1 13 Mar 2024 ran 4 of 7 samples (3 unverified)
Not all layers are equally as important: Every Layer Counts BERT 0 4 3 Nov 2023 not harvested
LM-CPPF: Paraphrasing-Guided Data Augmentation for Contrastive Prompt-Based Few-Shot Fine-Tuning 1 1 29 May 2023 not harvested
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions 1 5 27 Apr 2023 not harvested
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale 4 1 15 Aug 2022 ran 2 of 5 samples (3 unverified)
Adversarial Self-Attention for Language Understanding 1 2 25 Jun 2022 not harvested
Prune Once for All: Sparse Pre-Trained Language Models 2 9 10 Nov 2021 not harvested
Charformer: Fast Character Transformers via Gradient-based Subword Tokenization 2 1 23 Jun 2021 ran 7 of 10 samples (3 unverified)
Pay Attention to MLPs 20 1 17 May 2021 ran 34 of 44 samples (10 unverified; 11 pointer-only for licence)
FNet: Mixing Tokens with Fourier Transforms 12 2 9 May 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
How to Train BERT with an Academic Budget 4 1 15 Apr 2021 not harvested
RealFormer: Transformer Likes Residual Attention 5 1 21 Dec 2020 not harvested
A Statistical Framework for Low-bitwidth Training of Deep Neural Networks 2 1 27 Oct 2020 ran 1 of 4 samples (3 unverified; 1 pointer-only for licence)
Big Bird: Transformers for Longer Sequences 14 1 28 Jul 2020 ran 10 of 15 samples (5 unverified; 11 pointer-only for licence)
SqueezeBERT: What can computer vision teach NLP about efficient neural networks? 6 1 19 Jun 2020 ran 0 of 1 samples (1 unverified)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
What Do Questions Exactly Ask? MFAE: Duplicate Question Identification with Multi-Fusion Asking Emphasis 1 1 7 May 2020 not harvested
SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization 6 6 8 Nov 2019 ran 6 of 8 samples (2 unverified; 1 pointer-only for licence)
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 7 23 Oct 2019 ran 2 of 31 samples (29 unverified)
Q8BERT: Quantized 8Bit BERT 5 1 14 Oct 2019 ran 3 of 11 samples (8 unverified; 3 pointer-only for licence)
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations 48 1 26 Sep 2019 ran 46 of 126 samples (80 unverified; 22 pointer-only for licence)
TinyBERT: Distilling BERT for Natural Language Understanding 10 3 23 Sep 2019 ran 0 of 4 samples (4 unverified; 4 pointer-only for licence)
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT 0 1 12 Sep 2019 not harvested
StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding 0 1 13 Aug 2019 not harvested
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding 3 2 29 Jul 2019 ran 0 of 1 samples (1 unverified; 1 pointer-only for licence)
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 2 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
SpanBERT: Improving Pre-training by Representing and Predicting Spans 6 1 24 Jul 2019 ran 3 of 15 samples (12 unverified; 4 pointer-only for licence)
XLNet: Generalized Autoregressive Pretraining for Language Understanding 27 1 19 Jun 2019 ran 10 of 24 samples (14 unverified; 3 pointer-only for licence)
ERNIE: Enhanced Language Representation with Informative Entities 2 1 17 May 2019 ran 3 of 3 samples (0 unverified; 1 pointer-only for licence)

The full list of 40 is in the JSON twin.

Dataset loaders archive 2025-07-28

6 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (multiple, see the paper)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MultiNLI MisMatched dev
  • MultiNLI Matched dev
  • MultiNLI Dev MisMatched
  • MultiNLI Dev Matched
  • MNLI-mm
  • mnli_mismatched
  • MNLI-m
  • MNLI
  • MultiNLI-mismatched
  • MultiNLI-matched
  • multi_nli
  • MultiNLI Dev
  • MultiNLI

13 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections