Datasets › WebQuestions

WebQuestions

Introduced by Jonathan Berant et al. in Semantic Parsing on Freebase from Question-Answer Pairs1 Jan 2013 archive 2025-07-28

The WebQuestions dataset is a question answering dataset using Freebase as the knowledge base and contains 6,642 question-answer pairs. It was created by crawling questions through the Google Suggest API, and then obtaining answers using Amazon Mechanical Turk. The original split uses 3,778 examples for training and 2,032 for testing. All answers are defined as Freebase entities.

Example questions (answers) in the dataset include “Where did Edgar Allan Poe died?” (baltimore) or “What degrees did Barack Obama get?” (bachelor_of_arts, juris_doctor).

Source: Question Answering with Subgraph Embeddings Image Source: Berant et al

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

26 shown of 26 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 241. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Paths-over-Graph: Knowledge Graph Empowered Large Language Model Reasoning 1 1 18 Oct 2024 ran 1 of 12 samples (11 unverified)
Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models 1 9 26 Mar 2024 ran 8 of 10 samples (2 unverified)
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines 3 1 5 Oct 2023 ran 3 of 7 samples (4 unverified)
Tree of Thoughts: Deliberate Problem Solving with Large Language Models 6 1 17 May 2023 ran 7 of 24 samples (17 unverified)
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
FiDO: Fusion-in-Decoder optimized for stronger performance and faster inference 0 1 15 Dec 2022 not harvested
FiE: Building a Global Probability Space by Leveraging Early Fusion in Encoder for Open-Domain Question Answering 0 4 18 Nov 2022 not harvested
Measuring and Narrowing the Compositionality Gap in Language Models 1 1 7 Oct 2022 not harvested
ReAct: Synergizing Reasoning and Acting in Language Models 9 1 6 Oct 2022 ran 15 of 34 samples (19 unverified; 5 pointer-only for licence)
PaLM: Scaling Language Modeling with Pathways 7 3 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models 19 1 28 Jan 2022 ran 2 of 7 samples (5 unverified)
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts 0 1 13 Dec 2021 not harvested
JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs 1 4 19 Jun 2021 not harvested
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering 2 1 9 Jun 2021 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
UniK-QA: Unified Representations of Structured and Unstructured Knowledge for Open-Domain Question Answering 1 1 29 Dec 2020 not harvested
Language Models are Few-Shot Learners 67 4 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks 18 1 22 May 2020 ran 4 of 6 samples (2 unverified)
Toward Subgraph-Guided Knowledge Graph Question Generation with Graph Neural Networks 1 1 13 Apr 2020 ran 0 of 5 samples (5 unverified)
Dense Passage Retrieval for Open-Domain Question Answering 19 1 10 Apr 2020 ran 10 of 14 samples (4 unverified; 9 pointer-only for licence)
REALM: Retrieval-Augmented Language Model Pre-Training 6 1 10 Feb 2020 ran 4 of 4 samples (0 unverified)
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 1 23 Oct 2019 ran 2 of 31 samples (29 unverified)
Latent Retrieval for Weakly Supervised Open Domain Question Answering 3 1 1 Jun 2019 not harvested
Language Models are Unsupervised Multitask Learners 21 1 14 Feb 2019 not harvested
Large-scale Simple Question Answering with Memory Networks 3 1 5 Jun 2015 not harvested
Question Answering with Subgraph Embeddings 1 1 14 Jun 2014 not harvested
Open Question Answering with Weakly Supervised Embedding Models 0 1 16 Apr 2014 not harvested

Dataset loaders archive 2025-07-28

6 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • WebQuestions

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections