Datasets › OpenBookQA

OpenBookQA (OBQA)

Introduced by Todor Mihaylov et al. in Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering1 Jan 2018 archive 2025-07-28

OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject. It consists of 5,957 multiple-choice elementary-level science questions (4,957 train, 500 dev, 500 test), which probe the understanding of a small “book” of 1,326 core science facts and the application of these facts to novel situations. For training, the dataset includes a mapping from each question to the core science fact it was designed to probe. Answering OpenBookQA questions requires additional broad common knowledge, not contained in the book. The questions, by design, are answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. Additionally, the dataset includes a collection of 5,167 crowd-sourced common knowledge facts, and an expanded version of the train/dev/test questions where each question is associated with its originating core fact, a human accuracy score, a clarity score, and an anonymized crowd-worker ID.

Source: https://allenai.org/data/open-book-qa Image Source: https://arxiv.org/pdf/1809.02789.pdf

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering OpenBookQA GPT-4 + knowledge base Accuracy 95.9 — — 45 Compare
Question Answering OBQA FLAN 137B (zero-shot) Accuracy 78.4 Finetuned Language Models Are Zero-Shot Learners hiyouga/llama-efficient-tuning +7 9 Compare
Text Generation OpenBookQA no rows — — 0 Compare

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 635. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Mixture-of-Subspaces in Low-Rank Adaptation 1 1 16 Jun 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts 2 3 22 Apr 2024 ran 6 of 11 samples (5 unverified)
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions 1 6 27 Apr 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
GrapeQA: GRaph Augmentation and Pruning to Enhance Question-Answering 0 3 22 Mar 2023 not harvested
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
Large Language Models Can Self-Improve 0 6 20 Oct 2022 not harvested
Clues Before Answers: Generation-Enhanced Multiple-Choice QA 1 1 30 Apr 2022 ran 5 of 11 samples (6 unverified)
PaLM: Scaling Language Modeling with Pathways 7 2 5 Apr 2022 ran 30 of 37 samples (7 unverified)
GNN is a Counter? Revisiting GNN for Question Answering 0 1 7 Oct 2021 not harvested
Finetuned Language Models Are Zero-Shot Learners 8 2 3 Sep 2021 ran 0 of 1 samples (1 unverified)
QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering 6 3 13 Apr 2021 ran 1 of 25 samples (24 unverified; 4 pointer-only for licence)
Fusing Context Into Knowledge Graph for Commonsense Question Answering 2 2 9 Dec 2020 not harvested
Language Models are Few-Shot Learners 67 2 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
UnifiedQA: Crossing Format Boundaries With a Single QA System 2 1 2 May 2020 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
Careful Selection of Knowledge to solve Open Book Question Answering 0 1 24 Jul 2019 not harvested
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering 1 4 8 Sep 2018 ran 0 of 3 samples (3 unverified)

Dataset loaders archive 2025-07-28

7 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • OpenBookQA
  • OBQA

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections