Browse › Natural Language Processing › Natural Language Inference › WNLI

Natural Language Inference archive 2025-07-28

WNLI Benchmark (Natural Language Inference)

23 rows 21 with code listed 1 metric Dataset page

Natural language inference (NLI) is the task of determining whether a "hypothesis" is true (entailment), false (contradiction), or undetermined (neutral) given a "premise".

Example:

Premise Label Hypothesis
A man inspects the uniform of a figure in some East Asian country. contradiction The man is sleeping.
An older and younger man smiling. neutral Two men are smiling and laughing at the cats playing on the floor.
A soccer game with multiple males playing. entailment Some men are playing a sport.

Approaches used for NLI include earlier symbolic and statistical approaches to more recent deep learning approaches. Benchmark datasets used for NLI include SNLI, MultiNLI, SciTail, among others. You can get hands-on practice on the SNLI task by following this d2l.ai chapter.

Further readings:

The archive carries no text for this table; the description above is the archive's text for the task Natural Language Inference. archive 2025-07-28

Over time archive 2025-07-28

The chart needs JavaScript; the table below carries every value.

Direction inferred from the metric name, not from the archive: Accuracy (higher is better). Points are placed at the row's paper date; 22 of 23 rows carry one.

Results archive 2025-07-28

Archive rows end at the archive snapshot, 2025-07-28: no result published after that date is in this table. Rank is the archive's row order at that snapshot; not re-ranked here. Metric values are the archive's strings. Column headers sort the table in your browser; each row keeps its archive rank.

Paper Code Ran Syntology Report
1 Turing NLR v5 XXL 5.4B (fine-tuned) 95.9 – – – not matched report
2 DeBERTa 94.5 – Paper Code 2020 4 of 13 ran · 9 unverified report
3 T5-XXL 11B 93.2 – Paper Code 2019 2 of 31 ran · 29 unverified report
4 XLNet 92.5 – Paper Code 2019 10 of 24 ran · 14 unverified report
5 ALBERT 91.8 – Paper Code 2019 46 of 126 ran · 80 unverified report
6 T5-XL 3B 89.7 – Paper Code 2019 2 of 31 ran · 29 unverified report
7 StructBERTRoBERTa ensemble 89.7 – Paper – 2019 no code linked report
8 HNNensemble 89 – Paper Code 2019 0 of 1 ran · 1 unverified report
9 RoBERTa (ensemble) 89 – Paper Code 2019 22 of 48 ran · 26 unverified report
10 T5-Large 770M 85.6 – Paper Code 2019 2 of 31 ran · 29 unverified report
11 HNN 83.6 – Paper Code 2019 0 of 1 ran · 1 unverified report
12 T5-Base 220M 78.8 – Paper Code 2019 2 of 31 ran · 29 unverified report
13 BERTwiki 340M (fine-tuned on WSCR) 74.7 – Paper Code 2019 linked, not harvested report
14 FLAN 137B (zero-shot) 74.6 – Paper Code 2021 0 of 1 ran · 1 unverified report
15 BERT-large 340M (fine-tuned on WSCR) 71.9 – Paper Code 2019 linked, not harvested report
16 BERT-base 110M (fine-tuned on WSCR) 70.5 – Paper Code 2019 linked, not harvested report
17 FLAN 137B (few-shot, k=4) 70.4 – Paper Code 2021 0 of 1 ran · 1 unverified report
18 T5-Small 60M 69.2 – Paper Code 2019 2 of 31 ran · 29 unverified report
19 ERNIE 2.0 Large 67.8 – Paper Code 2019 0 of 1 ran · 1 unverified report
20 SqueezeBERT 65.1 – Paper Code 2020 0 of 1 ran · 1 unverified report
21 BERT-large 340M 65.1 – Paper Code 2018 204 of 659 ran · 455 unverified report
22 RWKV-4-Raven-14B 49.3 – Paper Code 2023 3 of 11 ran · 8 unverified report
23 DistilBERT 66M 44.4 – Paper Code 2019 19 of 27 ran · 8 unverified report

All 23 rows shown. 22 link to a paper page on this site; 0 are marked as using additional training data in the archive. No GitHub stars are tracked; "Code" is the first repository the archive lists for the row. The archive carries no row tags, review links or community-submitted rows for this table; none are shown. archive 2025-07-28

Syntology Ran reads "N of M ran · U unverified": of the M code samples Syntology harvested from repositories linked to that row's paper (joined by arXiv id), N executed on a synthesized input and the other U = M−N are unverified (harvested, no recorded run). It counts code from repositories linked to that row's paper, not this result: the row's number was not reproduced and nothing here is a correctness claim. The other cell texts mean no graph line for the row: "linked, not harvested" (the archive links code, Syntology has not harvested it), "no code linked" (no code link in the archive), "not matched" (the row's paper URL matched no paper on this site). 18 rows have a graph line, from 12 distinct papers; 12 rows (8 papers) have at least one sample that ran. Counting each paper once: Syntology ran 310 of 943 samples; 633 unverified. Separately, 201 of those 943 are pointer-only (licence): the site points at that code rather than redistributing it, a licence property recorded for ran and unverified samples alike; each cell's tooltip carries the row's own pointer-only count. Read from the graph 2026-09-24. Per-sample status is on the paper page.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections