Home › Datasets › task › Medical Diagnosis
Medical Diagnosis datasets
archive 2025-07-28
21 datasets carry the task tag "Medical Diagnosis" (the task itself: Medical Diagnosis), ordered by the archive's paper count. Page 1 of 1: 21 shown of 21. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Medical Diagnosis datasets 1–21 of 21
PadChest is a labeled large-scale, high resolution chest x-ray dataset for the automated exploration of medical images along with their associated reports.
116 papers · 0 benchmarks
IntrA is an open-access 3D intracranial aneurysm dataset that makes the application of points-based and mesh-based classification and segmentation models available.
27 papers · 2 benchmarks
BCI (Breast Cancer Immunohistochemical Image Generation)
The evaluation of human epidermal growth factor receptor 2 (HER2) expression is essential to formulate a precise treatment for breast cancer.
19 papers · 1 benchmark
DDXPlus (DDXPlus: A New Dataset For Automatic Medical Diagnosis)
There has been a rapidly growing interest in Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the machine learning research literature, aiming to assist doctors in telemedicine services.
19 papers · 0 benchmarks
MedConceptsQA - Open Source Medical Concepts QA Benchmark The benchmark can be found here: https://huggingface.co/datasets/ofir408/MedConceptsQA
13 papers · 2 benchmarks
REFLACX (Reports and eye-tracking data for localization of abnormalities in chest x-rays)
The REFLACX dataset contains eye-tracking data for 3,032 readings of chest x-rays by five radiologists.
10 papers · 0 benchmarks
LIMUC (Labeled Images for Ulcerative Colitis)
The LIMUC dataset is the largest publicly available labeled ulcerative colitis dataset that compromises 11276 images from 564 patients and 1043 colonoscopy procedures.
4 papers · 1 benchmark
The Kvasir-VQA dataset is an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations.
3 papers · 0 benchmarks
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
2 papers · 0 benchmarks
This dataset is created from MIMIC-III (Medical Information Mart for Intensive Care III) and contains simulated patient admission notes.
2 papers · 4 benchmarks
Background: Lung cancer risk classification is an increasingly important area of research as low-dose thoracic CT screening programs have become standard of care for patients at high risk for lung cancer.
2 papers · 1 benchmark
FracAtlas (A Dataset for Fracture Classification, Localization and Segmentation of Musculoskeletal Radiographs)
FractureAtlas is a musculoskeletal bone fracture dataset with annotations for deep learning tasks like classification, localization, and segmentation.
2 papers · 0 benchmarks
BreastDICOM4 ([MIMBCD-UI] UTA4: Medical Imaging DICOM Files Dataset)
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 1 benchmark
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 0 benchmarks
EBHI-Seg is a dataset containing 5,170 images of six types of tumor differentiation stages and the corresponding ground truth images.
1 paper · 0 benchmarks
A dataset for medical consultation dialogues.
1 paper · 0 benchmarks
HTDM (Hypertention Disease Medication)
Hypertention Disease Medication dataset.
1 paper · 0 benchmarks
LLM Health Benchmarks Dataset The Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties.
1 paper · 0 benchmarks
Liver-US (Liver Ultrasound Dataset for Medical Image Classification)
The Liver-US dataset is a comprehensive collection of high-quality ultrasound images of the liver, including both normal and abnormal cases.
1 paper · 1 benchmark
Onchocerciasis is causing blindness in over half a million people in the world today.
1 paper · 0 benchmarks
Sakha-TB (400+400 CXR images for TB diagnosis)
Sakha-TB is a de-identified image dataset of frontal chest X-rays (CXR), collected through collaboration with several medical institutions in the Republic of Sakha (Yakutia, Russia).
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.