Datasets › PubMedQA corpus with metadata

PubMedQA corpus with metadata

Introduced by Kunal Sawarkar et al. in MetaGen Blended RAG: Higher Accuracy for Domain-Specific Q&A Without Fine-Tuning23 May 2025 archive 2025-07-28

PubMedQA-MetaGen: Metadata-Enriched PubMedQA Corpus

Dataset Summary PubMedQA-MetaGen is a metadata-enriched version of the PubMedQA biomedical question-answering dataset, created using the MetaGenBlendedRAG enrichment pipeline. The dataset contains both the original and enriched versions of the corpus, enabling direct benchmarking of retrieval-augmented and semantic search approaches in biomedical NLP.

Files Provided PubMedQA_original_corpus.json This file contains the original PubMedQA corpus, formatted directly from the official PubMedQA dataset. Each record includes the biomedical question, context (abstract), and answer fields, mirroring the original dataset structure.

PubMedQA_corpus_with_metadata.json This file contains the metadata-enriched version, created by processing the original corpus through the MetaGenBlendedRAG pipeline. In addition to the original fields, each entry is augmented with structured metadata—including key concepts, MeSH terms, automatically generated keywords, extracted entities, and LLM-generated summaries—designed to support advanced retrieval and RAG research.

How to Use RAG evaluation: Benchmark your retrieval-augmented QA models using the enriched context for higher recall and precision. Semantic Search: Build improved biomedical search engines leveraging topic, entity, and keyword metadata. NLP & LLM Fine-tuning: Use for fine-tuning models that benefit from structured biomedical context. Dataset Structure Each sample contains:

Original fields: Question, context (abstract), answer

Enriched fields (in PubMedQA_corpus_with_metadata.json only):

Key concepts and topics Extracted MeSH terms and UMLS entities Automatically generated keywords Section/type labels LLM-generated summaries/metadata Document identifiers and links Dataset Creation Process Source: Original PubMedQA dataset. Metadata Enrichment: Applied the MetaGenBlendedRAG pipeline (rule-based, NLP, and LLM-driven enrichment). Outputs: Two files—original and enriched—supporting both traditional and metadata-driven research. Intended Use and Limitations For research and educational use in biomedical QA, RAG, semantic retrieval, and metadata enrichment evaluation. Note: Some metadata fields generated by LLMs may vary in quality; users should verify outputs for critical applications.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
RAG PubMedQA corpus with metadata MetaGen Blended RAG ANS-EM 77.90 MetaGen Blended RAG: Higher Accuracy for Domain-Specific... ibm-self-serve-assets/metagen-blended-rag 1 Compare
Retrieval PubMedQA corpus with metadata MetaGen Blended RAG Accuracy (Top-1) 82.1 MetaGen Blended RAG: Higher Accuracy for Domain-Specific... ibm-self-serve-assets/metagen-blended-rag 1 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MetaGen Blended RAG: Higher Accuracy for Domain-Specific Q&A Without Fine-Tuning 1 2 23 May 2025 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • PubMedQA corpus with metadata

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections