Datasets › BoolQ

BoolQ (Boolean Questions)

Introduced by Christopher Clark et al. in BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions1 Jan 2019 archive 2025-07-28

BoolQ is a question answering dataset for yes/no questions containing 15942 examples. These questions are naturally occurring – they are generated in unprompted and unconstrained settings. Each example is a triplet of (question, passage, answer), with the title of the page as optional additional context.

Questions are gathered from anonymized, aggregated queries to the Google search engine. Queries that are likely to be yes/no questions are heuristically identified and questions are only kept if a Wikipedia page is returned as one of the first five results, in which case the question and Wikipedia page are given to a human annotator for further processing. Annotators label question/article pairs in a three-step process. First, they decide if the question is good, meaning it is comprehensible, unambiguous, and requesting factual information. This judgment is made before the annotator sees the Wikipedia page. Next, for good questions, annotators find a passage within the document that contains enough information to answer the question. Annotators can mark questions as “not answerable” if the Wikipedia article does not contain the requested information. Finally, annotators mark whether the question’s answer is “yes” or “no”. Only questions that were marked as having a yes/no answer are used, and each question is paired with the selected passage instead of the entire document.

Source: https://github.com/google-research-datasets/boolean-questions Image Source: BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Benchmarks archive 2025-07-28

All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering BoolQ Mistral-Nemo 12B (HPT) Accuracy 99.87 Hierarchical Prompting Taxonomy: A Universal Evaluation... devichand579/HPT 65 Compare
parameter-efficient fine-tuning BoolQ LLaMA2-7b Accuracy (% ) 82.63 QLoRA: Efficient Finetuning of Quantized LLMs qwenlm/qwen +19 4 Compare
Classification BoolQ OPT-1.3B Test Accuracy 62.5% Achieving Dimension-Free Communication in Federated... ZidongLiu/DeComFL 2 Compare
Text Classification BoolQ no rows — — 0 Compare
Text Generation BoolQ no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 32 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 701. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments 0 1 15 Oct 2024 not harvested
GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs 1 1 27 Aug 2024 ran 3 of 5 samples (2 unverified; 3 pointer-only for licence)
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles 1 1 18 Jun 2024 not harvested
Mixture-of-Subspaces in Low-Rank Adaptation 1 1 16 Jun 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
Achieving Dimension-Free Communication in Federated Learning via Zeroth-Order Optimization 1 2 24 May 2024 ran 1 of 1 samples (0 unverified)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts 2 3 22 Apr 2024 ran 6 of 11 samples (5 unverified)
DoRA: Weight-Decomposed Low-Rank Adaptation 5 1 14 Feb 2024 ran 8 of 15 samples (7 unverified; 14 pointer-only for licence)
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 4 18 Jul 2023 ran 31 of 52 samples (21 unverified; 16 pointer-only for licence)
QLoRA: Efficient Finetuning of Quantized LLMs 20 1 23 May 2023 ran 17 of 26 samples (9 unverified; 17 pointer-only for licence)
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
Hyena Hierarchy: Towards Larger Convolutional Language Models 7 1 21 Feb 2023 ran 5 of 5 samples (0 unverified; 4 pointer-only for licence)
Hungry Hungry Hippos: Towards Language Modeling with State Space Models 3 5 28 Dec 2022 ran 7 of 15 samples (8 unverified)
OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization 1 6 22 Dec 2022 not harvested
Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE 0 2 4 Dec 2022 not harvested
Ask Me Anything: A simple strategy for prompting language models 3 3 5 Oct 2022 ran 2 of 2 samples (0 unverified)
AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model 1 1 2 Aug 2022 ran 1 of 1 samples (0 unverified)
N-Grammer: Augmenting Transformers with latent n-grams 2 1 13 Jul 2022 ran 0 of 6 samples (6 unverified)
UL2: Unifying Language Learning Paradigms 2 2 10 May 2022 ran 0 of 16 samples (16 unverified)
PaLM: Scaling Language Modeling with Pathways 7 1 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)
ST-MoE: Designing Stable and Transferable Sparse Expert Models 3 2 17 Feb 2022 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
Scaling Language Models: Methods, Analysis & Insights from Training Gopher 3 1 8 Dec 2021 not harvested
Finetuned Language Models Are Zero-Shot Learners 8 3 3 Sep 2021 ran 0 of 1 samples (1 unverified)
LoRA: Low-Rank Adaptation of Large Language Models 74 1 17 Jun 2021 ran 34 of 84 samples (50 unverified; 28 pointer-only for licence)
Entailment as Few-Shot Learner 3 1 29 Apr 2021 ran 1 of 3 samples (2 unverified)
Muppet: Massive Multi-task Representations with Pre-Finetuning 2 2 26 Jan 2021 not harvested
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
Language Models are Few-Shot Learners 67 2 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)

The full list of 32 is in the JSON twin.

Dataset loaders archive 2025-07-28

7 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA 3.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • BoolQ

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections