Datasets › BoolQ
BoolQ (Boolean Questions)
BoolQ is a question answering dataset for yes/no questions containing 15942 examples. These questions are naturally occurring – they are generated in unprompted and unconstrained settings. Each example is a triplet of (question, passage, answer), with the title of the page as optional additional context.
Questions are gathered from anonymized, aggregated queries to the Google search engine. Queries that are likely to be yes/no questions are heuristically identified and questions are only kept if a Wikipedia page is returned as one of the first five results, in which case the question and Wikipedia page are given to a human annotator for further processing. Annotators label question/article pairs in a three-step process. First, they decide if the question is good, meaning it is comprehensible, unambiguous, and requesting factual information. This judgment is made before the annotator sees the Wikipedia page. Next, for good questions, annotators find a passage within the document that contains enough information to answer the question. Annotators can mark questions as “not answerable” if the Wikipedia article does not contain the requested information. Finally, annotators mark whether the question’s answer is “yes” or “no”. Only questions that were marked as having a yes/no answer are used, and each question is paired with the selected passage instead of the entire document.
Source: https://github.com/google-research-datasets/boolean-questions Image Source: BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Benchmarks archive 2025-07-28
All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Question Answering | BoolQ | Mistral-Nemo 12B (HPT) Accuracy 99.87 | Hierarchical Prompting Taxonomy: A Universal Evaluation... | devichand579/HPT | 65 | Compare |
| parameter-efficient fine-tuning | BoolQ | LLaMA2-7b Accuracy (% ) 82.63 | QLoRA: Efficient Finetuning of Quantized LLMs | qwenlm/qwen +19 | 4 | Compare |
| Classification | BoolQ | OPT-1.3B Test Accuracy 62.5% | Achieving Dimension-Free Communication in Federated... | ZidongLiu/DeComFL | 2 | Compare |
| Text Classification | BoolQ | no rows | — | — | 0 | Compare |
| Text Generation | BoolQ | no rows | — | — | 0 | Compare |
Papers archive 2025-07-28
30 shown of 32 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 701. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
The full list of 32 is in the JSON twin.
Dataset loaders archive 2025-07-28
7 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- BoolQ
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections