Datasets › PIQA

PIQA (Physical Interaction: Question Answering)

Introduced by Yonatan Bisk et al. in PIQA: Reasoning about Physical Commonsense in Natural Language26 Nov 2019 archive 2025-07-28

PIQA is a dataset for commonsense reasoning, and was created to investigate the physical knowledge of existing models in NLP.

Source: PIQA

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering PIQA Unicorn 11B (fine-tuned) Accuracy 90.1 UNICORN on RAINBOW: A Universal Commonsense Reasoning... allenai/rainbow 67 Compare
Text Generation PIQA no rows — — 0 Compare

Papers archive 2025-07-28

28 shown of 28 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 772. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments 0 1 15 Oct 2024 not harvested
Mixture-of-Subspaces in Low-Rank Adaptation 1 1 16 Jun 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts 2 3 22 Apr 2024 ran 6 of 11 samples (5 unverified)
Mixtral of Experts 6 2 8 Jan 2024 ran 5 of 5 samples (0 unverified)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks 2 1 5 Jan 2024 not harvested
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning 2 3 10 Oct 2023 ran 3 of 3 samples (0 unverified)
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
Textbooks Are All You Need II: phi-1.5 technical report 1 1 11 Sep 2023 not harvested
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 4 18 Jul 2023 ran 31 of 52 samples (21 unverified; 16 pointer-only for licence)
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions 1 6 27 Apr 2023 not harvested
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling 4 4 3 Apr 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot 6 5 2 Jan 2023 ran 2 of 12 samples (10 unverified; 9 pointer-only for licence)
Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question Answering 1 2 29 Oct 2022 ran 1 of 7 samples (6 unverified)
Task Compass: Scaling Multi-task Pre-training with Task Prefix 1 3 12 Oct 2022 not harvested
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)
Efficient Language Modeling with Sparse all-MLP 0 4 14 Mar 2022 not harvested
Scaling Language Models: Methods, Analysis & Insights from Training Gopher 3 1 8 Dec 2021 not harvested
Finetuned Language Models Are Zero-Shot Learners 8 2 3 Sep 2021 ran 0 of 1 samples (1 unverified)
UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark 1 1 24 Mar 2021 ran 0 of 5 samples (5 unverified)
Language Models are Few-Shot Learners 67 2 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
UnifiedQA: Crossing Format Boundaries With a Single QA System 2 1 2 May 2020 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
PIQA: Reasoning about Physical Commonsense in Natural Language 2 4 26 Nov 2019 not harvested
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism 10 1 17 Sep 2019 ran 12 of 47 samples (35 unverified; 15 pointer-only for licence)
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 534 1 11 Oct 2018 ran 204 of 659 samples (455 unverified; 149 pointer-only for licence)

Dataset loaders archive 2025-07-28

7 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • PIQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections