Browse › Natural Language Processing › Natural Language Inference › SNLI
SNLI Benchmark (Natural Language Inference)
Natural language inference (NLI) is the task of determining whether a "hypothesis" is true (entailment), false (contradiction), or undetermined (neutral) given a "premise".
Example:
| Premise | Label | Hypothesis |
|---|---|---|
| A man inspects the uniform of a figure in some East Asian country. | contradiction | The man is sleeping. |
| An older and younger man smiling. | neutral | Two men are smiling and laughing at the cats playing on the floor. |
| A soccer game with multiple males playing. | entailment | Some men are playing a sport. |
Approaches used for NLI include earlier symbolic and statistical approaches to more recent deep learning approaches. Benchmark datasets used for NLI include SNLI, MultiNLI, SciTail, among others. You can get hands-on practice on the SNLI task by following this d2l.ai chapter.
Further readings:
The archive carries no text for this table; the description above is the archive's text for the task Natural Language Inference. archive 2025-07-28
Over time archive 2025-07-28
The chart needs JavaScript; the table below carries every value.
Direction inferred from the metric name, not from the archive: % Test Accuracy (higher is better), % Train Accuracy (higher is better), Dev Accuracy (higher is better), % Dev Accuracy (higher is better), Accuracy (higher is better). Not inferred (points only, no best-so-far line): Parameters. Points are placed at the row's paper date; 94 of 98 rows carry one.
Results archive 2025-07-28
Archive rows end at the archive snapshot, 2025-07-28: no result published after that date is in this table. Rank is the archive's row order at that snapshot; not re-ranked here. Metric values are the archive's strings. Column headers sort the table in your browser; each row keeps its archive rank.
| Paper | Code | Ran Syntology | Report | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | UnitedSynT5 (3B) | 94.7 | ✓ | Paper | – | 2024 | no code linked | report | |||||
| 2 | UnitedSynT5 (335M) | 93.5 | ✓ | Paper | – | 2024 | no code linked | report | |||||
| 3 | Neural Tree Indexers for Text Understanding | 93.1 | 355 | – | Paper | Code | 2021 | 1 of 3 ran · 2 unverified | report | ||||
| 4 | EFL (Entailment as Few-shot Learner) + RoBERTa-large | 93.1 | ? | 355m | – | Paper | Code | 2021 | 1 of 3 ran · 2 unverified | report | |||
| 5 | RoBERTa-large+Self-Explaining | 92.3 | 340 | – | Paper | Code | 2020 | linked, not harvested | report | ||||
| 6 | RoBERTa-large + self-explaining layer | 92.3 | ? | 355m+ | – | Paper | Code | 2020 | linked, not harvested | report | |||
| 7 | CA-MTL | 92.1 | 92.6 | 340m | – | Paper | Code | 2020 | linked, not harvested | report | |||
| 8 | SemBERT | 91.9 | 94.4 | 339m | – | Paper | Code | 2019 | 4 of 12 ran · 8 unverified | report | |||
| 9 | MT-DNN-SMARTLARGEv0 | 91.7 | 92.6 | – | Paper | Code | 2019 | 8 of 8 ran · 0 unverified | report | ||||
| 10 | MT-DNN | 91.6 | 97.2 | 330m | – | Paper | Code | 2019 | 7 of 13 ran · 6 unverified | report | |||
| 11 | SJRC (BERT-Large +SRL) | 91.3 | 95.7 | 308m | – | Paper | – | 2018 | no code linked | report | |||
| 12 | Ntumpha | 90.5 | 99.1 | 220 | – | Paper | Code | 2019 | 7 of 13 ran · 6 unverified | report | |||
| 13 | Densely-Connected Recurrent and Co-Attentive Network Ensemble | 90.1 | 95.0 | 53.3m | – | Paper | – | 2018 | no code linked | report | |||
| 14 | MFAE | 90.07 | 93.18 | – | Paper | Code | 2020 | linked, not harvested | report | ||||
| 15 | Fine-Tuned LM-Pretrained Transformer | 89.9 | 96.6 | 85m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 16 | 300D DMAN Ensemble | 89.6 | 96.1 | 79m | – | Paper | Code | 2019 | linked, not harvested | report | |||
| 17 | 300D DMAN Ensemble | 89.6 | 96.1 | 79m | – | – | – | not matched | report | ||||
| 18 | 150D Multiway Attention Network Ensemble | 89.4 | 95.5 | 58m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 19 | 450D DR-BiLSTM Ensemble | 89.3 | 94.8 | 45m | – | Paper | – | 2018 | no code linked | report | |||
| 20 | 300D CAFE Ensemble | 89.3 | 92.5 | 17.5m | – | Paper | – | 2017 | no code linked | report | |||
| 21 | ESIM + ELMo Ensemble | 89.3 | 92.1 | 40m | – | Paper | Code | 2018 | 24 of 58 ran · 34 unverified | report | |||
| 22 | KIM Ensemble | 89.1 | 93.6 | 43m | – | Paper | Code | 2017 | linked, not harvested | report | |||
| 23 | SLRC | 89.1 | 89.1 | 6.1m | – | Paper | – | 2018 | no code linked | report | |||
| 24 | RE2 | 88.9 | 94.0 | 2.8m | – | Paper | Code | 2019 | 1 of 1 ran · 0 unverified | report | |||
| 25 | Densely-Connected Recurrent and Co-Attentive Network | 88.9 | 93.1 | 6.7m | – | Paper | – | 2018 | no code linked | report | |||
| 26 | DEIM | 88.9 | 92.6 | 22m | – | Paper | – | 2022 | no code linked | report | |||
| 27 | 448D Densely Interactive Inference Network (DIIN, code) Ensemble | 88.9 | 92.3 | 17m | – | Paper | Code | 2017 | 1 of 1 ran · 0 unverified | report | |||
| 28 | 300D DMAN | 88.8 | 95.4 | 9.2m | – | Paper | Code | 2019 | linked, not harvested | report | |||
| 29 | 300D DMAN | 88.8 | 95.4 | 9.2m | – | – | – | not matched | report | ||||
| 30 | BiMPM Ensemble | 88.8 | 93.2 | 6.4m | – | Paper | Code | 2017 | 1 of 4 ran · 3 unverified | report | |||
| 31 | ESIM + ELMo | 88.7 | 91.6 | 8.0m | – | Paper | Code | 2018 | 24 of 58 ran · 34 unverified | report | |||
| 32 | KIM | 88.6 | 94.1 | 4.3m | – | Paper | Code | 2017 | linked, not harvested | report | |||
| 33 | 600D ESIM + 300D Syntactic TreeLSTM | 88.6 | 93.5 | 7.7m | – | Paper | Code | 2016 | linked, not harvested | report | |||
| 34 | 450D DR-BiLSTM | 88.5 | 94.1 | 7.5m | – | Paper | – | 2018 | no code linked | report | |||
| 35 | Stochastic Answer Network | 88.5 | 93.3 | 3.5m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 36 | 300D CAFE | 88.5 | 89.8 | 4.7m | – | Paper | – | 2017 | no code linked | report | |||
| 37 | 150D Multiway Attention Network | 88.3 | 94.5 | 14m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 38 | Biattentive Classification Network + CoVe + Char | 88.1 | 88.5 | 22m | – | Paper | Code | 2017 | 3 of 3 ran · 0 unverified | report | |||
| 39 | aESIM | 88.1 | – | Paper | – | 2018 | no code linked | report | |||||
| 40 | 448D Densely Interactive Inference Network (DIIN, code) | 88.0 | 91.2 | 4.4m | – | Paper | Code | 2017 | 1 of 1 ran · 0 unverified | report | |||
| 41 | Enhanced Sequential Inference Model (Chen et al., [2017a]) | 88.0 | – | Paper | Code | 2016 | linked, not harvested | report | |||||
| 42 | BiMPM | 87.5 | 90.9 | 1.6m | – | Paper | Code | 2017 | 1 of 4 ran · 3 unverified | report | |||
| 43 | 300D re-read LSTM | 87.5 | 90.7 | 2.0m | – | – | – | not matched | report | ||||
| 44 | 300D re-read LSTM | 87.5 | 90.7 | 2.0m | – | Paper | – | 2016 | no code linked | report | |||
| 45 | 2400D Multiple-Dynamic Self-Attention Model | 87.4 | 89.0 | 7.0m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 46 | 300D Full tree matching NTI-SLSTM-LSTM w/ global attention | 87.3 | 88.5 | 3.2m | – | Paper | Code | 2016 | linked, not harvested | report | |||
| 47 | 300D 2-layer Bi-CAS-LSTM | 87 | – | Paper | – | 2018 | no code linked | report | |||||
| 48 | 200D decomposable attention feed-forward model with intra-sentence attention | 86.8 | 90.5 | 580k | – | Paper | Code | 2016 | 0 of 6 ran · 6 unverified | report | |||
| 49 | 200D decomposable attention model with intra-sentence attention | 86.8 | 90.5 | 580k | – | Paper | Code | 2016 | 0 of 6 ran · 6 unverified | report | |||
| 50 | 600D Dynamic Self-Attention Model | 86.8 | 87.3 | 2.1m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 51 | CBS-1 + ESIM | 86.73 | – | Paper | – | 2018 | no code linked | report | |||||
| 52 | 512D Dynamic Meta-Embeddings | 86.7 | 91.6 | 9m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 53 | 600D BiLSTM with generalized pooling | 86.6 | 94.9 | 65m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 54 | 600D Hierarchical BiLSTM with Max Pooling (HBMP, code) | 86.6 | 89.9 | 22m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 55 | Densely-Connected Recurrent and Co-Attentive Network (encoder) | 86.5 | 91.4 | 5.6m | – | Paper | – | 2018 | no code linked | report | |||
| 56 | 300D Reinforced Self-Attention Network | 86.3 | 92.6 | 3.1m | – | Paper | Code | 2018 | linked, not harvested | report | |||
| 57 | Distance-based Self-Attention Network | 86.3 | 89.6 | 4.7m | – | Paper | – | 2017 | no code linked | report | |||
| 58 | 200D decomposable attention feed-forward model | 86.3 | 89.5 | 380k | – | Paper | Code | 2016 | 0 of 6 ran · 6 unverified | report | |||
| 59 | 200D decomposable attention model | 86.3 | 89.5 | 380k | – | Paper | Code | 2016 | 0 of 6 ran · 6 unverified | report | |||
| 60 | 450D LSTMN with deep attention fusion | 86.3 | 88.5 | 3.4m | – | Paper | Code | 2016 | 2 of 2 ran · 0 unverified | report | |||
| 61 | 300D mLSTM word-by-word attention model | 86.1 | 92.0 | 1.9m | – | Paper | Code | 2015 | 1 of 1 ran · 0 unverified | report | |||
| 62 | 600D Gumbel TreeLSTM encoders | 86.0 | 93.1 | 10m | – | Paper | Code | 2017 | linked, not harvested | report | |||
| 63 | 600D Residual stacked encoders | 86.0 | 91.0 | 29m | – | Paper | Code | 2017 | linked, not harvested | report | |||
| 64 | Star-Transformer (no cross sentence attention) | 86.0 | – | Paper | Code | 2019 | 0 of 1 ran · 1 unverified | report | |||||
| 65 | 300D CAFE (no cross-sentence attention) | 85.9 | 87.3 | 3.7m | – | Paper | – | 2017 | no code linked | report | |||
| 66 | 1200D REGMAPR (Base+Reg) | 85.9 | – | – | – | – | – | not matched | report | ||||
| 67 | 300D Residual stacked encoders | 85.7 | 89.8 | 9.7m | – | Paper | Code | 2017 | linked, not harvested | report | |||
| 68 | 300D LSTMN with deep attention fusion | 85.7 | 87.3 | 1.7m | – | Paper | Code | 2016 | 2 of 2 ran · 0 unverified | report | |||
| 69 | 300D Gumbel TreeLSTM encoders | 85.6 | 91.2 | 2.9m | – | Paper | Code | 2017 | linked, not harvested | report | |||
| 70 | 300D Directional self-attention network encoders | 85.6 | 91.1 | 2.4m | – | Paper | Code | 2017 | linked, not harvested | report | |||
| 71 | 600D (300+300) Deep Gated Attn. BiLSTM encoders | 85.5 | 90.5 | 12m | – | Paper | Code | 2017 | linked, not harvested | report | |||
| 72 | 300D MMA-NSE encoders with attention | 85.4 | 86.9 | 3.2m | – | Paper | Code | 2016 | linked, not harvested | report | |||
| 73 | 50D stacked TC-LSTMs | 85.1 | 86.7 | 190k | – | Paper | – | 2016 | no code linked | report | |||
| 74 | 600D (300+300) BiLSTM encoders with intra-attention and symbolic preproc. | 85.0 | 85.9 | 2.8m | – | Paper | Code | 2016 | linked, not harvested | report | |||
| 75 | Stacked Bi-LSTMs (shortcut connections, max-pooling) | 84.8 | – | Paper | Code | 2018 | linked, not harvested | report | |||||
| 76 | 300D NSE encoders | 84.6 | 86.2 | 3.0m | – | Paper | Code | 2016 | linked, not harvested | report | |||
| 77 | 100D DF-LSTM | 84.6 | 85.2 | 320k | – | Paper | – | 2016 | no code linked | report | |||
| 78 | 4096D BiLSTM with max-pooling | 84.5 | 85.6 | 40m | – | Paper | Code | 2017 | 6 of 7 ran · 1 unverified | report | |||
| 79 | Bi-LSTM sentence encoder (max-pooling) | 84.5 | – | Paper | Code | 2018 | linked, not harvested | report | |||||
| 80 | Stacked Bi-LSTMs (shortcut connections, max-pooling, attention) | 84.4 | – | Paper | Code | 2018 | linked, not harvested | report | |||||
| 81 | 600D (300+300) BiLSTM encoders with intra-attention | 84.2 | 84.5 | 2.8m | – | Paper | Code | 2016 | linked, not harvested | report | |||
| 82 | SWEM-max | 83.8 | – | Paper | Code | 2018 | linked, not harvested | report | |||||
| 83 | 100D LSTMs w/ word-by-word attention | 83.5 | 85.3 | 250k | – | Paper | Code | 2015 | 1 of 1 ran · 0 unverified | report | |||
| 84 | 300D NTI-SLSTM-LSTM encoders | 83.4 | 82.5 | 4.0m | – | Paper | Code | 2016 | linked, not harvested | report | |||
| 85 | 600D (300+300) BiLSTM encoders | 83.3 | 86.4 | 2.0m | – | Paper | Code | 2016 | linked, not harvested | report | |||
| 86 | 300D SPINN-PI encoders | 83.2 | 89.2 | 3.7m | – | Paper | Code | 2016 | 0 of 6 ran · 6 unverified | report | |||
| 87 | 300D Tree-based CNN encoders | 82.1 | 83.3 | 3.5m | – | Paper | – | 2015 | no code linked | report | |||
| 88 | 1024D GRU encoders w/ unsupervised 'skip-thoughts' pre-training | 81.4 | 98.8 | 15m | – | Paper | Code | 2015 | linked, not harvested | report | |||
| 89 | DELTA (LSTM) | 80.7 | – | Paper | Code | 2019 | linked, not harvested | report | |||||
| 90 | 300D LSTM encoders | 80.6 | 83.9 | 3.0m | – | Paper | Code | 2016 | 0 of 6 ran · 6 unverified | report | |||
| 91 | + Unigram and bigram features | 78.2 | 99.7 | – | Paper | Code | 2015 | 3 of 3 ran · 0 unverified | report | ||||
| 92 | 100D LSTM encoders | 77.6 | 84.8 | 220k | – | Paper | Code | 2015 | 3 of 3 ran · 0 unverified | report | |||
| 93 | Unlexicalized features | 50.4 | 49.4 | – | Paper | Code | 2015 | 3 of 3 ran · 0 unverified | report | ||||
| 94 | MT-DNN-SMART_100%ofTrainingData | 91.6 | – | Paper | Code | 2019 | 8 of 8 ran · 0 unverified | report | |||||
| 95 | MT-DNN-SMART_10%ofTrainingData | 88.7 | – | Paper | Code | 2019 | 8 of 8 ran · 0 unverified | report | |||||
| 96 | MT-DNN-SMART_1%ofTrainingData | 86 | – | Paper | Code | 2019 | 8 of 8 ran · 0 unverified | report | |||||
| 97 | MT-DNN-SMART_0.1%ofTrainingData | 82.7 | – | Paper | Code | 2019 | 8 of 8 ran · 0 unverified | report | |||||
| 98 | SplitEE-S | 79.0 | – | Paper | Code | 2023 | linked, not harvested | report |
All 98 rows shown. 94 link to a paper page on this site; 2 are marked as using additional training data in the archive. No GitHub stars are tracked; "Code" is the first repository the archive lists for the row. The archive carries no row tags, review links or community-submitted rows for this table; none are shown. archive 2025-07-28
Syntology Ran reads "N of M ran · U unverified": of the M code samples Syntology harvested from repositories linked to that row's paper (joined by arXiv id), N executed on a synthesized input and the other U = M−N are unverified (harvested, no recorded run). It counts code from repositories linked to that row's paper, not this result: the row's number was not reproduced and nothing here is a correctness claim. The other cell texts mean no graph line for the row: "linked, not harvested" (the archive links code, Syntology has not harvested it), "no code linked" (no code link in the archive), "not matched" (the row's paper URL matched no paper on this site). 33 rows have a graph line, from 17 distinct papers; 26 rows (14 papers) have at least one sample that ran. Counting each paper once: Syntology ran 63 of 130 samples; 67 unverified. Separately, 45 of those 130 are pointer-only (licence): the site points at that code rather than redistributing it, a licence property recorded for ran and unverified samples alike; each cell's tooltip carries the row's own pointer-only count. Read from the graph 2026-09-25. Per-sample status is on the paper page.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections