Datasets › ARC (AI2 Reasoning Challenge)
ARC (AI2 Reasoning Challenge)
The AI2’s Reasoning Challenge (ARC) dataset is a multiple-choice question-answering dataset, containing questions from science exams from grade 3 to grade 9. The dataset is split in two partitions: Easy and Challenge, where the latter partition contains the more difficult questions that require reasoning. Most of the questions have 4 answer choices, with <1% of all the questions having either 3 or 5 answer choices. ARC includes a supporting KB of 14.3M unstructured text passages.
Source: Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question Answering Image Source: https://arxiv.org/abs/1803.05457
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Common Sense Reasoning | ARC (Challenge) | GPT-4 (few-shot, k=25) Accuracy 96.4 | GPT-4 Technical Report | openai/evals +10 | 54 | Compare |
| Common Sense Reasoning | ARC (Easy) | ST-MoE-32B 269B (fine-tuned) Accuracy 95.2 | ST-MoE: Designing Stable and Transferable Sparse Expert Models | tensorflow/mesh +2 | 47 | Compare |
| Stance Detection | ARC (AI2 Reasoning Challenge) | TESTED F1 64.82 | Topic-Guided Sampling For Data-Efficient Multi-Domain... | copenlu/TESTED +1 | 1 | Compare |
Papers archive 2025-07-28
23 shown of 23 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 178. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Mixture-of-Subspaces in Low-Rank Adaptation | 1 | 2 | 16 Jun 2024 | ran 4 of 6 samples (2 unverified; 6 pointer-only for licence) |
| MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts | 2 | 6 | 22 Apr 2024 | ran 6 of 11 samples (5 unverified) |
| Mixtral of Experts | 6 | 2 | 8 Jan 2024 | ran 5 of 5 samples (0 unverified) |
| Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks | 2 | 2 | 5 Jan 2024 | not harvested |
| Mamba: Linear-Time Sequence Modeling with Selective State Spaces | 35 | 1 | 1 Dec 2023 | ran 18 of 62 samples (44 unverified; 28 pointer-only for licence) |
| Mistral 7B | 6 | 2 | 10 Oct 2023 | ran 9 of 11 samples (2 unverified; 1 pointer-only for licence) |
| Textbooks Are All You Need II: phi-1.5 technical report | 1 | 2 | 11 Sep 2023 | not harvested |
| Model Card and Evaluations for Claude Models | 0 | 3 | 11 Jul 2023 | not harvested |
| Stay on topic with Classifier-Free Guidance | 0 | 4 | 30 Jun 2023 | not harvested |
| Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection | 2 | 1 | 1 Jun 2023 | not harvested |
| PaLM 2 Technical Report | 1 | 7 | 17 May 2023 | not harvested |
| Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling | 4 | 4 | 3 Apr 2023 | not harvested |
| BloombergGPT: A Large Language Model for Finance | 2 | 8 | 30 Mar 2023 | not harvested |
| GPT-4 Technical Report | 11 | 2 | 15 Mar 2023 | ran 2 of 5 samples (3 unverified; 1 pointer-only for licence) |
| LLaMA: Open and Efficient Foundation Language Models | 57 | 8 | 27 Feb 2023 | ran 26 of 58 samples (32 unverified; 4 pointer-only for licence) |
| SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot | 6 | 10 | 2 Jan 2023 | ran 2 of 12 samples (10 unverified; 9 pointer-only for licence) |
| Galactica: A Large Language Model for Science | 1 | 8 | 16 Nov 2022 | ran 0 of 2 samples (2 unverified) |
| Large Language Models Can Self-Improve | 0 | 6 | 20 Oct 2022 | not harvested |
| UL2: Unifying Language Learning Paradigms | 2 | 6 | 10 May 2022 | ran 0 of 16 samples (16 unverified) |
| ST-MoE: Designing Stable and Transferable Sparse Expert Models | 3 | 4 | 17 Feb 2022 | ran 5 of 5 samples (0 unverified; 5 pointer-only for licence) |
| GLaM: Efficient Scaling of Language Models with Mixture-of-Experts | 0 | 4 | 13 Dec 2021 | not harvested |
| Finetuned Language Models Are Zero-Shot Learners | 8 | 4 | 3 Sep 2021 | ran 0 of 1 samples (1 unverified) |
| Language Models are Few-Shot Learners | 67 | 4 | 28 May 2020 | ran 15 of 65 samples (50 unverified; 4 pointer-only for licence) |
Dataset loaders archive 2025-07-28
4 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- ARC (Easy)
- ARC (Challenge)
- ARC (AI2 Reasoning Challenge)
3 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections