Datasets › ARC (AI2 Reasoning Challenge)

ARC (AI2 Reasoning Challenge)

Introduced by Peter Clark et al. in Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge1 Jan 2018 archive 2025-07-28

The AI2’s Reasoning Challenge (ARC) dataset is a multiple-choice question-answering dataset, containing questions from science exams from grade 3 to grade 9. The dataset is split in two partitions: Easy and Challenge, where the latter partition contains the more difficult questions that require reasoning. Most of the questions have 4 answer choices, with <1% of all the questions having either 3 or 5 answer choices. ARC includes a supporting KB of 14.3M unstructured text passages.

Source: Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question Answering Image Source: https://arxiv.org/abs/1803.05457

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Common Sense Reasoning ARC (Challenge) GPT-4 (few-shot, k=25) Accuracy 96.4 GPT-4 Technical Report openai/evals +10 54 Compare
Common Sense Reasoning ARC (Easy) ST-MoE-32B 269B (fine-tuned) Accuracy 95.2 ST-MoE: Designing Stable and Transferable Sparse Expert Models tensorflow/mesh +2 47 Compare
Stance Detection ARC (AI2 Reasoning Challenge) TESTED F1 64.82 Topic-Guided Sampling For Data-Efficient Multi-Domain... copenlu/TESTED +1 1 Compare

Papers archive 2025-07-28

23 shown of 23 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 178. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Mixture-of-Subspaces in Low-Rank Adaptation 1 2 16 Jun 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts 2 6 22 Apr 2024 ran 6 of 11 samples (5 unverified)
Mixtral of Experts 6 2 8 Jan 2024 ran 5 of 5 samples (0 unverified)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks 2 2 5 Jan 2024 not harvested
Mamba: Linear-Time Sequence Modeling with Selective State Spaces 35 1 1 Dec 2023 ran 18 of 62 samples (44 unverified; 28 pointer-only for licence)
Mistral 7B 6 2 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
Textbooks Are All You Need II: phi-1.5 technical report 1 2 11 Sep 2023 not harvested
Model Card and Evaluations for Claude Models 0 3 11 Jul 2023 not harvested
Stay on topic with Classifier-Free Guidance 0 4 30 Jun 2023 not harvested
Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection 2 1 1 Jun 2023 not harvested
PaLM 2 Technical Report 1 7 17 May 2023 not harvested
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling 4 4 3 Apr 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 8 30 Mar 2023 not harvested
GPT-4 Technical Report 11 2 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
LLaMA: Open and Efficient Foundation Language Models 57 8 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot 6 10 2 Jan 2023 ran 2 of 12 samples (10 unverified; 9 pointer-only for licence)
Galactica: A Large Language Model for Science 1 8 16 Nov 2022 ran 0 of 2 samples (2 unverified)
Large Language Models Can Self-Improve 0 6 20 Oct 2022 not harvested
UL2: Unifying Language Learning Paradigms 2 6 10 May 2022 ran 0 of 16 samples (16 unverified)
ST-MoE: Designing Stable and Transferable Sparse Expert Models 3 4 17 Feb 2022 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts 0 4 13 Dec 2021 not harvested
Finetuned Language Models Are Zero-Shot Learners 8 4 3 Sep 2021 ran 0 of 1 samples (1 unverified)
Language Models are Few-Shot Learners 67 4 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • ARC (Easy)
  • ARC (Challenge)
  • ARC (AI2 Reasoning Challenge)

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections