Datasets › COPA

COPA (Choice of Plausible Alternatives)

Introduced in Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning1 Jan 2011 archive 2025-07-28

The Choice Of Plausible Alternatives (COPA) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning. COPA consists of 1000 questions, split equally into development and test sets of 500 questions each. Each question is composed of a premise and two alternatives, where the task is to select the alternative that more plausibly has a causal relation with the premise. The correct alternative is randomized so that the expected performance of randomly guessing is 50%.

Source: Choice of Plausible Alternatives (COPA) Image Source: https://people.ict.usc.edu/~gordon/copa.html

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering COPA PaLM 540B (finetuned) Accuracy 100 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 60 Compare

Papers archive 2025-07-28

23 shown of 23 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 329. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning 2 1 23 May 2023 not harvested
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
Exploring the Benefits of Training Expert Language Models over Instruction Tuning 2 1 7 Feb 2023 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Hungry Hungry Hippos: Towards Language Modeling with State Space Models 3 5 28 Dec 2022 ran 7 of 15 samples (8 unverified)
Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE 0 2 4 Dec 2022 not harvested
Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models 0 1 28 Oct 2022 not harvested
Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners 1 1 6 Oct 2022 not harvested
Ask Me Anything: A simple strategy for prompting language models 3 3 5 Oct 2022 ran 2 of 2 samples (0 unverified)
AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model 1 1 2 Aug 2022 ran 1 of 1 samples (0 unverified)
N-Grammer: Augmenting Transformers with latent n-grams 2 1 13 Jul 2022 ran 0 of 6 samples (6 unverified)
UL2: Unifying Language Learning Paradigms 2 2 10 May 2022 ran 0 of 16 samples (16 unverified)
PaLM: Scaling Language Modeling with Pathways 7 1 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Efficient Language Modeling with Sparse all-MLP 0 5 14 Mar 2022 not harvested
ST-MoE: Designing Stable and Transferable Sparse Expert Models 3 2 17 Feb 2022 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
KELM: Knowledge Enhanced Pre-Trained Language Representations with Message Passing on Hierarchical Relational Graphs 1 1 9 Sep 2021 ran 1 of 6 samples (5 unverified)
Finetuned Language Models Are Zero-Shot Learners 8 3 3 Sep 2021 ran 0 of 1 samples (1 unverified)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 2 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
Language Models are Few-Shot Learners 67 5 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 4 23 Oct 2019 ran 2 of 31 samples (29 unverified)
WinoGrande: An Adversarial Winograd Schema Challenge at Scale 10 5 24 Jul 2019 not harvested
SocialIQA: Commonsense Reasoning about Social Interactions 1 2 22 Apr 2019 ran 0 of 10 samples (10 unverified)
Handling Multiword Expressions in Causality Estimation 0 5 1 Jan 2017 not harvested

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

BSD 2-Clause License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • COPA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections