Browse › Natural Language Processing › Question Answering › PIQA

Question Answering archive 2025-07-28

PIQA Benchmark (Question Answering)

67 rows 62 with code listed 1 metric Dataset page

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

The archive carries no text for this table; the description above is the archive's text for the task Question Answering. archive 2025-07-28

Over time archive 2025-07-28

The chart needs JavaScript; the table below carries every value.

Direction inferred from the metric name, not from the archive: Accuracy (higher is better). Points are placed at the row's paper date; 67 of 67 rows carry one.

Results archive 2025-07-28

Archive rows end at the archive snapshot, 2025-07-28: no result published after that date is in this table. Rank is the archive's row order at that snapshot; not re-ranked here. Metric values are the archive's strings. Column headers sort the table in your browser; each row keeps its archive rank.

Paper Code Ran Syntology Report
1 Unicorn 11B (fine-tuned) 90.1 – Paper Code 2021 3 of 5 ran · 2 unverified report
2 LLaMA3 8B+MoSLoRA 89.7 – Paper Code 2024 4 of 6 ran · 2 unverified report
3 CompassMTL 567M with Tailor 88.3 – Paper Code 2022 linked, not harvested report
4 LLaMA-3 8B + MixLoRA 87.6 – Paper Code 2024 6 of 11 ran · 5 unverified report
5 DeBERTa-Large 304M 87.4 – Paper Code 2022 2 of 7 ran · 5 unverified report
6 CompassMTL 567M 87.3 – Paper Code 2022 linked, not harvested report
7 LLaMA-2 13B + MixLoRA 86.8 – Paper Code 2024 6 of 11 ran · 5 unverified report
8 Shakti-LLM (2.5B) 86.2 – Paper – 2024 no code linked report
9 DeBERTa-Large 304M (classification-based) 85.9 – Paper Code 2022 2 of 7 ran · 5 unverified report
10 ExDeBERTa 567M 85.5 – Paper Code 2022 linked, not harvested report
11 UnifiedQA 3B 85.3 – Paper Code 2020 4 of 7 ran · 3 unverified report
12 PaLM 2-L (1-shot) 85.0 – Paper Code 2023 linked, not harvested report
13 Mixtral 8x7B (0-shot) 83.6 – Paper Code 2024 5 of 5 ran · 0 unverified report
14 PaLM 2-M (1-shot) 83.2 – Paper Code 2023 linked, not harvested report
15 LLaMA-2 7B + MixLoRA 83.2 – Paper Code 2024 6 of 11 ran · 5 unverified report
16 Mistral 7B (0-shot) 83.0 – Paper Code 2023 10 of 11 ran · 1 unverified report
17 LLaMA 65B (0-shot) 82.8 – Paper Code 2023 37 of 58 ran · 21 unverified report
18 LLaMA 2 70B (0-shot) 82.8 – Paper Code 2023 31 of 52 ran · 21 unverified report
19 Camelidae-8×34B 82.7 – Paper Code 2024 linked, not harvested report
20 LLaMA 33B (0-shot) 82.3 – Paper Code 2023 37 of 58 ran · 21 unverified report
21 PaLM 2-S (1-shot) 82.2 – Paper Code 2023 linked, not harvested report
22 Mistral 7B (0-shot) 82.2 – Paper Code 2024 5 of 5 ran · 0 unverified report
23 MT-NLG 530B (0-shot) 82.0 – Paper Code 2019 12 of 47 ran · 35 unverified report
24 LLaMA 2 34B (0-shot) 81.9 – Paper Code 2023 31 of 52 ran · 21 unverified report
25 Gopher 280B (0-shot) 81.8 – Paper Code 2021 linked, not harvested report
26 Chinchilla 70B (0-shot) 81.8 – Paper Code 2022 8 of 11 ran · 3 unverified report
27 FLAN 137B (few-shot, k=10) 81.7 – Paper Code 2021 0 of 1 ran · 1 unverified report
28 OPT-175B 81.07 – Paper Code 2023 9 of 12 ran · 3 unverified report
29 GPT-3 175B (0-shot) 81.0 – Paper Code 2020 41 of 65 ran · 24 unverified report
30 SparseGPT 175B (50% Sparsity) 80.63 – Paper Code 2023 9 of 12 ran · 3 unverified report
31 FLAN 137B (0-shot) 80.5 – Paper Code 2021 0 of 1 ran · 1 unverified report
32 LLaMA 2 13B (0-shot) 80.5 – Paper Code 2023 31 of 52 ran · 21 unverified report
33 LLaMA 13B (0-shot) 80.1 – Paper Code 2023 37 of 58 ran · 21 unverified report
34 LLaMA 7B (0-shot) 79.8 – Paper Code 2023 37 of 58 ran · 21 unverified report
35 SparseGPT 175B (4:8 Sparsity) 79.54 – Paper Code 2023 9 of 12 ran · 3 unverified report
36 SparseGPT 175B (2:4 Sparsity) 79.54 – Paper Code 2023 9 of 12 ran · 3 unverified report
37 RoBERTa-Large 355M 79.4 – Paper Code 2019 37 of 48 ran · 11 unverified report
38 LLaMA 2 7B (0-shot) 78.8 – Paper Code 2023 31 of 52 ran · 21 unverified report
39 Bloomberg GPT 50B (1-shot) 77.9 – Paper Code 2023 linked, not harvested report
40 OPT 66B (1-shot) 77.6 – Paper Code 2023 linked, not harvested report
41 RoBERTa-large 355M (fine-tuned) 77.1 – Paper Code 2019 linked, not harvested report
42 phi-1.5-web (1.3B) 77 – Paper Code 2023 linked, not harvested report
43 BLOOM 176B (1-shot) 77 – Paper Code 2023 linked, not harvested report
44 Pythia 12B (5-shot) 76.7 – Paper Code 2023 linked, not harvested report
45 Open-LLaMA-3B-v2 76.2 – Paper Code 2023 3 of 3 ran · 0 unverified report
46 Pythia 12B (0-shot) 76 – Paper Code 2023 linked, not harvested report
47 Sheared-LLaMA-2.7B 75.8 – Paper Code 2023 3 of 3 ran · 0 unverified report
48 GPT-NeoX 20B (1-shot) 75.8 – Paper Code 2023 linked, not harvested report
49 Pythia 6.9B (0-shot) 75.2 – Paper Code 2023 linked, not harvested report
50 Sheared-LLaMA-1.3B 73.4 – Paper Code 2023 3 of 3 ran · 0 unverified report
51 sMLP - deterministic 9.4B (0-shot) 73 – Paper – 2022 no code linked report
52 GPT-3 Large 760M (0-shot) 72.9 – Paper Code 2020 41 of 65 ran · 24 unverified report
53 FLAN-T5-Large 783M 72.2 – Paper Code 2023 linked, not harvested report
54 LaMini-GPT 1.5B 71.3 – Paper Code 2023 linked, not harvested report
55 LaMini-F-T5 783M 70.6 – Paper Code 2023 linked, not harvested report
56 GPT-2-XL 1.5B 70.5 – Paper Code 2023 linked, not harvested report
57 Pythia 1B (5-shot) 70.4 – Paper Code 2023 linked, not harvested report
58 GPT-2-small 124M (fine-tuned) 69.2 – Paper Code 2019 linked, not harvested report
59 Gshard 9B 68.1 – Paper – 2022 no code linked report
60 LaMini-T5 738M 67.2 – Paper Code 2023 linked, not harvested report
61 BERT-large 340M (fine-tuned) 66.8 – Paper Code 2019 linked, not harvested report
62 BERT-Large 340M 66.7 – Paper Code 2018 208 of 659 ran · 451 unverified report
63 Base Layers 10B (0-shot) 63.8 – Paper – 2022 no code linked report
64 HASH Layers 10B (0-shot) 63.8 – Paper – 2022 no code linked report
65 T5-Large 738M 55.9 – Paper Code 2023 linked, not harvested report
66 OPT-175B (50% Sparsity) 54.73 – Paper Code 2023 9 of 12 ran · 3 unverified report
67 Random chance baseline 50 – Paper Code 2019 linked, not harvested report

All 67 rows shown. 67 link to a paper page on this site; 0 are marked as using additional training data in the archive. No GitHub stars are tracked; "Code" is the first repository the archive lists for the row. The archive carries no row tags, review links or community-submitted rows for this table; none are shown. archive 2025-07-28

Syntology Ran reads "N of M ran · U unverified": of the M code samples Syntology harvested from repositories linked to that row's paper (joined by arXiv id), N executed on a synthesized input and the other U = M−N are unverified (harvested, no recorded run). It counts code from repositories linked to that row's paper, not this result: the row's number was not reproduced and nothing here is a correctness claim. The other cell texts mean no graph line for the row: "linked, not harvested" (the archive links code, Syntology has not harvested it), "no code linked" (no code link in the archive), "not matched" (the row's paper URL matched no paper on this site). 35 rows have a graph line, from 17 distinct papers; 33 rows (16 papers) have at least one sample that ran. Counting each paper once: Syntology ran 420 of 1,008 samples; 588 unverified. Separately, 232 of those 1,008 are pointer-only (licence): the site points at that code rather than redistributing it, a licence property recorded for ran and unverified samples alike; each cell's tooltip carries the row's own pointer-only count. Read from the graph 2026-09-25. Per-sample status is on the paper page.

Since the archive: results placed by Syntology Syntology

Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the rows above. Measured before publishing: the extractor was run on 883 held-out archive papers and processed 881 of them (the other 2 failed with an error before producing any output and are not part of this measurement); on those 881 papers, a blind reviewer judged 108 of 110 accepted entries correct (95% lower confidence bound 0.9361). Syntology has checked 6,885 of the 9,623 papers on this site that are newer than the archive; results from the others appear after they are checked.

Paper Method (configuration) Accuracy Date Where in the paper Code Syntology Report
Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax... arXiv:2609.18131 2.54 Colla-Q 80.5 16 Sep 2026 Table 1, row “2.54 Colla-Q” repository linked, no samples harvested report
SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference arXiv:2606.26587 Llama-3.1-8B SharQ 80.20 25 Jun 2026 Table 1, row “Llama-3.1-8B SharQ” 2 of 3 ran report
HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM... arXiv:2605.29843 4 HARP · Llama 2 7B 78.3 28 May 2026 Table 2, row “4 HARP” 8 of 9 ran report
YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference arXiv:2604.13556 YOCO++ 71.16 15 Apr 2026 Table 1, row “YOCO++” 3 of 5 ran report

Syntology 4 entries, one per paper, newest first by month (the arXiv date, else the month in the arXiv id), then by arXiv id. Each value is the cell text as the paper prints it; hover "Where in the paper" for the table's caption and each value's column header. Not part of the archive and not in the chart above. Drawn from 6,885 of 9,623 newer papers checked so far.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections