Home › Datasets › modality › Medical
Medical datasets
archive 2025-07-28
394 datasets carry the modality tag "Medical", ordered by the archive's paper count. Page 1 of 9: 48 shown of 394. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Medical datasets 1–48 of 394
MIMIC-III (The Medical Information Mart for Intensive Care III)
The Medical Information Mart for Intensive Care III (MIMIC-III) dataset is a large, de-identified and publicly-available collection of medical records.
1,041 papers · 8 benchmarks
The CheXpert dataset contains 224,316 chest radiographs of 65,240 patients with both frontal and lateral views available.
628 papers · 3 benchmarks
The fastMRI dataset includes two types of MRI scans: knee MRIs and the brain (neuro) MRIs, and containing training, validation, and masked test sets.
332 papers · 5 benchmarks
DRIVE (Digital Retinal Images for Vessel Extraction)
The Digital Retinal Images for Vessel Extraction (DRIVE) dataset is a dataset for retinal vessel segmentation.
311 papers · 2 benchmarks
The LIDC-IDRI dataset contains lesion annotations from four experienced thoracic radiologists.
240 papers · 6 benchmarks
MIMIC-CXR from Massachusetts Institute of Technology presents 371,920 chest X-rays associated with 227,943 imaging studies from 65,079 patients.
240 papers · 3 benchmarks
ChestX-ray14 is a medical imaging dataset which comprises 112,120 frontal-view X-ray images of 30,805 (collected from the year of 1992 to 2015) unique patients with the text-mined fourteen common disease labels, mined from the text…
237 papers · 6 benchmarks
HAM10000 is a dataset of 10000 training images for detecting pigmented skin lesions.
209 papers · 2 benchmarks
Kvasir-SEG is an open-access dataset of gastrointestinal polyp images and corresponding segmentation masks, manually annotated by a medical doctor and then verified by an experienced gastroenterologist.
201 papers · 2 benchmarks
CAMELYON16 (Cancer Metastases in Lymph Nodes Challenge 2016)
The dataset consists of 400 whole-slide images (WSIs) of lymph node sections stained with hematoxylin and eosin (H&E), collected from two medical centers in the Netherlands.
172 papers · 1 benchmark
CORD-19 is a free resource of tens of thousands of scholarly articles about COVID-19, SARS-CoV-2, and related coronaviruses for use by the global research community.
163 papers · 1 benchmark
STARE (Structured Analysis of the Retina)
The STARE (Structured Analysis of the Retina) dataset is a dataset for retinal vessel segmentation.
146 papers · 6 benchmarks
VQA-RAD (Visual Question Answering in Radiology)
VQA-RAD consists of 3,515 question–answer pairs on 315 radiology images.
145 papers · 0 benchmarks
Cholec80 is an endoscopic video dataset containing 80 videos of cholecystectomy surgeries performed by 13 surgeons.
134 papers · 2 benchmarks
The LUNA challenges provide datasets for automatic nodule detection algorithms using the largest publicly available reference database of chest CT scans, the LIDC-IDRI data set.
125 papers · 2 benchmarks
The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
PadChest is a labeled large-scale, high resolution chest x-ray dataset for the automated exploration of medical images along with their associated reports.
116 papers · 0 benchmarks
GlaS (Gland Segmentation in Colon Histology Images Challenge)
The dataset used in this challenge consists of 165 images derived from 16 H&E stained histological sections of stage T3 or T42 colorectal adenocarcinoma.
112 papers · 1 benchmark
PatchCamelyon is an image classification dataset.
110 papers · 4 benchmarks
Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse…
105 papers · 1 benchmark
The LUNA16 (LUng Nodule Analysis) dataset is a dataset for lung segmentation.
99 papers · 0 benchmarks
The Medical Segmentation Decathlon is a collection of medical image segmentation datasets.
97 papers · 1 benchmark
The sleep-edf database contains 197 whole-night PolySomnoGraphic sleep recordings, containing EEG, EOG, chin EMG, and event markers.
94 papers · 5 benchmarks
PathVQA consists of 32,799 open-ended questions from 4,998 pathology images where each question is manually checked to ensure correctness.
89 papers · 0 benchmarks
PPMI (Parkinson’s Progression Markers Initiative)
The Parkinson’s Progression Markers Initiative (PPMI) dataset originates from an observational clinical and longitudinal study comprising evaluations of people with Parkinson’s disease (PD), those people with high risk, and those who are…
87 papers · 3 benchmarks
CAMUS (Cardiac Acquisitions for Multi-structure Ultrasound Segmentation)
This project aims to provide all the materials to the community to resolve the problem of echocardiographic image segmentation and volume estimation from 2D ultrasound sequences (both two and four-chamber views).
85 papers · 0 benchmarks
The PROMISE12 dataset was made available for the MICCAI 2012 prostate segmentation challenge.
84 papers · 2 benchmarks
ChestX-ray8 is a medical imaging dataset which comprises 108,948 frontal-view X-ray images of 32,717 (collected from the year of 1992 to 2015) unique patients with the text-mined eight common disease labels, mined from the text…
81 papers · 0 benchmarks
RadGraph (RadGraph: Extracting Clinical Entities and Relations from Radiology Reports)
RadGraph is a dataset of entities and relations in radiology reports based on our novel information extraction schema, consisting of 600 reports with 30K radiologist annotations and 221K reports with 10.5M automatically generated…
78 papers · 0 benchmarks
The BraTS 2015 dataset is a dataset for brain tumor image segmentation.
69 papers · 1 benchmark
UBFC-rPPG (Univ. Bourgogne Franche-Comté Remote PhotoPlethysmoGraphy)
We introduce here a new database called UBFC-rPPG (stands for Univ.
69 papers · 1 benchmark
CoNSeP (Colorectal Nuclear Segmentation and Phenotypes)
The colorectal nuclear segmentation and phenotypes (CoNSeP) dataset consists of 41 H&E stained image tiles, each of size 1,000×1,000 pixels at 40× objective magnification.
68 papers · 2 benchmarks
SLAKE is an English-Chinese bilingual dataset consisting of 642 images and 14,028 question-answer pairs for training and testing Med-VQA systems.
63 papers · 0 benchmarks
PanNuke is a semi automatically generated nuclei instance segmentation and classification dataset with exhaustive nuclei labels across 19 different tissue types.
61 papers · 4 benchmarks
CHASEDB1 is a dataset for retinal vessel segmentation which contains 28 color retina images with the size of 999×960 pixels which are collected from both left and right eyes of 14 school children.
59 papers · 2 benchmarks
PMC-VQA is a large-scale medical visual question-answering dataset that contains 227k VQA pairs of 149k images that cover various modalities or diseases.
55 papers · 2 benchmarks
CVC-ClinicDB is an open-access dataset of 612 images with a resolution of 384×288 from 31 colonoscopy sequences.It is used for medical image segmentation, in particular polyp detection in colonoscopy videos.
54 papers · 1 benchmark
HRF (High-Resolution Fundus)
The HRF dataset is a dataset for retinal vessel segmentation which comprises 45 images and is organized as 15 subsets.
53 papers · 3 benchmarks
MedMentions is a new manually annotated resource for the recognition of biomedical concepts.
48 papers · 1 benchmark
- We present a large and diverse abdominal CT organ segmentation dataset, termed AbdomenCT-1K, with more than 1000 (1K) CT scans from 12 medical centers, including multi-phase, multi-vendor, and multi-disease cases.
46 papers · 0 benchmarks
LiTS17 (Liver Tumor Segmentation Challenge 2017)
LiTS17 is a liver tumor segmentation benchmark.
45 papers · 3 benchmarks
A large dataset of musculoskeletal radiographs containing 40,561 images from 14,863 studies, where each study is manually labeled by radiologists as either normal or abnormal.
43 papers · 0 benchmarks
VerSe (Large Scale Vertebrae Segmentation Challenge)
Spine or vertebral segmentation is a crucial step in all applications regarding automated quantification of spinal morphology and pathology.
39 papers · 0 benchmarks
BIOSSES (Biomedical Semantic Similarity Estimation System)
The BIOSSES data set comprises total 100 sentence pairs all of which were selected from the "TAC2 Biomedical Summarization Track Training Data Set" .
38 papers · 2 benchmarks
WORD (Whole abdominal Organs Dataset)
WORD is a dataset for organ semantic segmentation that contains 150 abdominal CT volumes (30,495 slices) and each volume has 16 organs with fine pixel-level annotations and scribble-based sparse annotation, which may be the largest dataset…
38 papers · 0 benchmarks
BRATS 2013 is a brain tumor segmentation dataset consists of synthetic and real images, where each of them is further divided into high-grade gliomas (HG) and low-grade gliomas (LG).
36 papers · 2 benchmarks
Contains hundreds of frontal view X-rays and is the largest public resource for COVID-19 image and prognostic data, making it a necessary resource to develop and evaluate tools to aid in the treatment of COVID-19.
35 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.