Datasets › MedConceptsQA

MedConceptsQA

Introduced by Ofir Ben Shoham et al. in MedConceptsQA: Open Source Medical Concepts QA Benchmark12 May 2024 archive 2025-07-28

MedConceptsQA - Open Source Medical Concepts QA Benchmark

The benchmark can be found here: https://huggingface.co/datasets/ofir408/MedConceptsQA

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Zero-Shot Learning MedConceptsQA gpt-4-0125-preview Accuracy 52.489 GPT-4 Technical Report openai/evals +10 13 Compare
Few-Shot Learning MedConceptsQA gpt-4-0125-preview Accuracy 61.911 GPT-4 Technical Report openai/evals +10 12 Compare

Papers archive 2025-07-28

12 shown of 12 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 13. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MedConceptsQA: Open Source Medical Concepts QA Benchmark 1 1 12 May 2024 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks 0 2 30 Mar 2024 not harvested
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains 1 2 15 Feb 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
MEDITRON-70B: Scaling Medical Pretraining for Large Language Models 1 4 27 Nov 2023 ran 9 of 14 samples (5 unverified)
Zephyr: Direct Distillation of LM Alignment 2 2 25 Oct 2023 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine 1 2 18 Aug 2023 not harvested
GPT-4 Technical Report 11 2 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
LLaMA: Open and Efficient Foundation Language Models 57 2 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records 0 1 2 Feb 2022 not harvested
Clinical-Longformer and Clinical-BigBird: Transformers for long clinical sequences 1 2 27 Jan 2022 not harvested
Language Models are Few-Shot Learners 67 2 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
BioBERT: a pre-trained biomedical language representation model for biomedical text mining 19 2 25 Jan 2019 ran 4 of 25 samples (21 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MedConceptsQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections