{"url":"/dataset/pubmedqa-corpus-with-metadata","name":"PubMedQA corpus with metadata","full_name":null,"description_markdown":"PubMedQA-MetaGen: Metadata-Enriched PubMedQA Corpus\r\n\r\nDataset Summary\r\nPubMedQA-MetaGen is a metadata-enriched version of the PubMedQA biomedical question-answering dataset, created using the MetaGenBlendedRAG enrichment pipeline. The dataset contains both the original and enriched versions of the corpus, enabling direct benchmarking of retrieval-augmented and semantic search approaches in biomedical NLP.\r\n\r\nFiles Provided\r\nPubMedQA_original_corpus.json This file contains the original PubMedQA corpus, formatted directly from the official PubMedQA dataset. Each record includes the biomedical question, context (abstract), and answer fields, mirroring the original dataset structure.\r\n\r\nPubMedQA_corpus_with_metadata.json This file contains the metadata-enriched version, created by processing the original corpus through the MetaGenBlendedRAG pipeline. In addition to the original fields, each entry is augmented with structured metadata—including key concepts, MeSH terms, automatically generated keywords, extracted entities, and LLM-generated summaries—designed to support advanced retrieval and RAG research.\r\n\r\nHow to Use\r\nRAG evaluation: Benchmark your retrieval-augmented QA models using the enriched context for higher recall and precision.\r\nSemantic Search: Build improved biomedical search engines leveraging topic, entity, and keyword metadata.\r\nNLP & LLM Fine-tuning: Use for fine-tuning models that benefit from structured biomedical context.\r\nDataset Structure\r\nEach sample contains:\r\n\r\nOriginal fields: Question, context (abstract), answer\r\n\r\nEnriched fields (in PubMedQA_corpus_with_metadata.json only):\r\n\r\nKey concepts and topics\r\nExtracted MeSH terms and UMLS entities\r\nAutomatically generated keywords\r\nSection/type labels\r\nLLM-generated summaries/metadata\r\nDocument identifiers and links\r\nDataset Creation Process\r\nSource: Original PubMedQA dataset.\r\nMetadata Enrichment: Applied the MetaGenBlendedRAG pipeline (rule-based, NLP, and LLM-driven enrichment).\r\nOutputs: Two files—original and enriched—supporting both traditional and metadata-driven research.\r\nIntended Use and Limitations\r\nFor research and educational use in biomedical QA, RAG, semantic retrieval, and metadata enrichment evaluation.\r\nNote: Some metadata fields generated by LLMs may vary in quality; users should verify outputs for critical applications.","description_withheld":null,"homepage":"https://huggingface.co/datasets/Shivam6693/PubMedQA-MetaGenBlendedRAG","introduced_date":"2025-05-23","introduced_date_note":null,"introduced_by":{"paper":"/paper/metagen-blended-rag-higher-accuracy-for","title":"MetaGen Blended RAG: Higher Accuracy for Domain-Specific Q&A Without Fine-Tuning","first_author":"Kunal Sawarkar","url":null},"license":{"name":"CC BY 4.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Retrieval","url":"/task/retrieval","datasets_with_task":"/datasets/task/retrieval"},{"name":"RAG","url":"/task/rag","datasets_with_task":"/datasets/task/rag"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["PubMedQA corpus with metadata"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/rag-on-pubmedqa-corpus-with-metadata","task":"RAG","dataset_variant":"PubMedQA corpus with metadata","rows":1,"metrics":["ANS-EM"],"first_row_in_archive_order":{"model":"MetaGen Blended RAG","paper":"/paper/metagen-blended-rag-higher-accuracy-for","metrics":{"ANS-EM":"77.90"},"code_links":[{"title":"ibm-self-serve-assets/metagen-blended-rag","url":"https://github.com/ibm-self-serve-assets/metagen-blended-rag"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/retrieval-on-pubmedqa-corpus-with-metadata","task":"Retrieval","dataset_variant":"PubMedQA corpus with metadata","rows":1,"metrics":["Accuracy (Top-1)"],"first_row_in_archive_order":{"model":"MetaGen Blended RAG","paper":"/paper/metagen-blended-rag-higher-accuracy-for","metrics":{"Accuracy (Top-1)":"82.1"},"code_links":[{"title":"ibm-self-serve-assets/metagen-blended-rag","url":"https://github.com/ibm-self-serve-assets/metagen-blended-rag"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/metagen-blended-rag-higher-accuracy-for","title":"MetaGen Blended RAG: Higher Accuracy for Domain-Specific Q&A Without Fine-Tuning","date":"2025-05-23","rows_on_this_dataset":2,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}