Datasets › PQAref
PQAref (Pubmed Question Answering with references)
The PQAref dataset is a dataset for fine-tuning large language models for referenced question-answering in biomedical domain.
The dataset contains 3 components:
Instruction - question that is supposed to be answered Abstracts - set of 10 relevant abstracts retrieved from PubMed by an IR system. They contain the PubMed id, abstract title and the content of the abstract Answer - expected answer, with references in the form of PubMed IDs.
The dataset was created semi-automatically, utilizing questions available from PubMedQA dataset.
The dataset contains 9,075 samples, split into training, validation and test set in proportion 80%:10%:10%.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
AGPLv3
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- PQAref
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections