Datasets › PubMedQA

PubMedQA

Introduced by Qiao Jin et al. in PubMedQA: A Dataset for Biomedical Research Question Answering13 Sep 2019 archive 2025-07-28

The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts.

PubMedQA has 1k expert labeled, 61.2k unlabeled and 211.3k artificially generated QA instances.

Source: PubMedQA Image Source: https://arxiv.org/pdf/1909.06146v1.pdf

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering PubMedQA Meditron-70B (CoT + SC) Accuracy 81.6 MEDITRON-70B: Scaling Medical Pretraining for Large... epfllm/meditron 30 Compare
Few-Shot Learning PubMedQA MetaGen Blended RAG (zero-shot) Accuracy 77.9 MetaGen Blended RAG: Higher Accuracy for Domain-Specific... ibm-self-serve-assets/metagen-blended-rag 2 Compare
Retrieval PubMedQA MetaGen Blended RAG Accuracy (Top-1) 82.1 MetaGen Blended RAG: Higher Accuracy for Domain-Specific... ibm-self-serve-assets/metagen-blended-rag 1 Compare

Papers archive 2025-07-28

17 shown of 17 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 276. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MetaGen Blended RAG: Higher Accuracy for Domain-Specific Q&A Without Fine-Tuning 1 3 23 May 2025 not harvested
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs 0 1 2 Jul 2024 not harvested
Evaluation of large language model performance on the Biomedical Language Understanding and Reasoning Benchmark 0 1 17 May 2024 not harvested
The Claude 3 Model Family: Opus, Sonnet, Haiku 0 2 4 Mar 2024 not harvested
MediSwift: Efficient Sparse Pre-trained Biomedical Language Models 0 1 1 Mar 2024 not harvested
MEDITRON-70B: Scaling Medical Pretraining for Large Language Models 1 1 27 Nov 2023 ran 9 of 14 samples (5 unverified)
BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine 1 1 18 Aug 2023 not harvested
The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning 2 2 23 May 2023 not harvested
Towards Expert-Level Medical Question Answering with Large Language Models 1 3 16 May 2023 not harvested
Large Language Models Encode Clinical Knowledge 1 7 26 Dec 2022 not harvested
Galactica: A Large Language Model for Science 1 3 16 Nov 2022 ran 0 of 2 samples (2 unverified)
BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining 4 2 19 Oct 2022 not harvested
Can large language models reason about medical questions? 1 1 17 Jul 2022 ran 0 of 10 samples (10 unverified)
LinkBERT: Pretraining Language Models with Document Links 1 2 29 Mar 2022 ran 0 of 14 samples (14 unverified)
BioELECTRA:Pretrained Biomedical text Encoder using Discriminators 1 1 11 Jun 2021 not harvested
Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing 2 1 31 Jul 2020 not harvested
PubMedQA: A Dataset for Biomedical Research Question Answering 5 1 13 Sep 2019 not harvested

Dataset loaders archive 2025-07-28

5 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • PubMedQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections