Home › Datasets › modality › Medical

Medical datasets

archive 2025-07-28

394 datasets carry the modality tag "Medical", ordered by the archive's paper count. Page 3 of 9: 48 shown of 394. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Medical datasets 97–144 of 394

GUE (Genome Understanding Evaluation)
A collection of $28$ datasets across $7$ tasks constructed for genome language model evaluation.
13 papers · 7 benchmarks
The ACNE04 dataset includes 3756 Chinese face images with Acne.
12 papers · 1 benchmark
ADAM (Adam: automatic detection challenge on age-related macular degeneration)
ADAM is organized as a half day Challenge, a Satellite Event of the ISBI 2020 conference in Iowa City, Iowa, USA.
12 papers · 1 benchmark
HyperKvasir dataset contains 110,079 images and 374 videos where it captures anatomical landmarks and pathological and normal findings.
12 papers · 2 benchmarks
4D-OR includes a total of 6734 scenes, recorded by six calibrated RGB-D Kinect sensors 1 mounted to the ceiling of the OR, with one frame-per-second, providing synchronized RGB and depth images.
11 papers · 3 benchmarks
ARCH is a computational pathology (CP) multiple instance captioning dataset to facilitate dense supervision of CP tasks.
11 papers · 0 benchmarks
BCN20000 is a dataset composed of 19,424 dermoscopic images of skin lesions captured from 2010 to 2016 in the facilities of the Hospital Clínic in Barcelona.
11 papers · 0 benchmarks
BreakHis (Breast Cancer Histopathological Database)
The Breast Cancer Histopathological Image Classification (BreakHis) is composed of 9,109 microscopic images of breast tumor tissue collected from 82 patients using different magnifying factors (40X, 100X, 200X, and 400X).
11 papers · 5 benchmarks
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
11 papers · 0 benchmarks
Kaggle EyePACS (Kaggle EyePACS. Diabetic Retinopathy Detection Identify signs of diabetic retinopathy in eye images)
Diabetic retinopathy is the leading cause of blindness in the working-age population of the developed world.
11 papers · 1 benchmark
The SD-198 dataset contains 198 different diseases from different types of eczema, acne and various cancerous conditions.
11 papers · 0 benchmarks
VinDr-CXR is an open large-scale dataset of chest X-rays with radiologist’s annotations.
11 papers · 0 benchmarks
CMeEE (Chinese Medical Named Entity Recognition Dataset)
Chinese Medical Named Entity Recognition, a dataset first released in CHIP20204, is used for CMeEE task.
10 papers · 1 benchmark
ChestX-Det is a chest X-Ray dataset with instance-level annotations (boxes and masks).
10 papers · 0 benchmarks
GRAZPEDWRI-DX is a public dataset of 20,327 pediatric wrist trauma X-ray images released by the University of Medicine of Graz.
10 papers · 4 benchmarks
The MM-WHS 2017 dataset is a dataset for multi-modality whole heart segmentation.
10 papers · 1 benchmark
OLIVES Dataset (Ophthalmic Labels for Investigating Visual Eye Semantics)
Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans.
10 papers · 0 benchmarks
REFLACX (Reports and eye-tracking data for localization of abnormalities in chest x-rays)
The REFLACX dataset contains eye-tracking data for 3,032 readings of chest x-rays by five radiologists.
10 papers · 0 benchmarks
TMED (Tufts Medical Echocardiogram Dataset)
TMED is a clinically-motivated benchmark dataset for computer vision and machine learning from limited labeled data.
10 papers · 0 benchmarks
2012 i2b2 Temporal Relations (2012 i2b2 Temporal Relations Corpus)
The Sixth Informatics for Integrating Biology and the Bedside (i2b2) Natural Language Processing Challenge for Clinical Records focused on the temporal relations in clinical narratives.
9 papers · 2 benchmarks
This dataset contains 1200 images (1000 WLI images and 200 FICE images) with fine-grained segmentation annotations.
9 papers · 1 benchmark
CHAOS (CHAOS - Combined (CT-MR) Healthy Abdominal Organ Segmentation)
CHAOS challenge aims the segmentation of abdominal organs (liver, kidneys and spleen) from CT and MRI data.
9 papers · 0 benchmarks
CTSpine1K is a large-scale and comprehensive dataset for research in spinal image analysis.
9 papers · 0 benchmarks
A large publicly available retinal fundus image dataset for glaucoma classification called G1020.
9 papers · 0 benchmarks
GLOBEM is a multi-year passive sensing datasets, containing over 700 user-years and 497 unique users' data collected from mobile and wearable sensors, together with a wide range of well-being metrics.
9 papers · 0 benchmarks
LoDoPaB-CT is a dataset of computed tomography images and simulated low-dose measurements.
9 papers · 1 benchmark
MedNLI (Medical Natural Language Inference)
The MedNLI dataset consists of the sentence pairs developed by Physicians from the Past Medical History section of MIMIC-III clinical notes annotated for Definitely True, Maybe True and Definitely False.
9 papers · 2 benchmarks
Mindboggle is a large publicly available dataset of manually labeled brain MRI.
9 papers · 0 benchmarks
PSI-AVA is a dataset designed for holistic surgical scene understanding.
9 papers · 0 benchmarks
RadQA (A Question Answering Dataset to Improve Comprehension of Radiology Reports)
RadQA is a radiology question answering dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans (resulting in a total of 6148 question-answer evidence pairs) by physicians.
9 papers · 1 benchmark
ATLAS v2.0 (Anatomical Tracings of Lesions After Stroke Dataset version 2.0)
Accurate lesion segmentation is critical in stroke rehabilitation research for the quantification of lesion burden and accurate image processing.
8 papers · 1 benchmark
Under a close collaboration with an expert radiologist team of the Hospital Universitario San Cecilio, the COVIDGR-1.0 dataset of patients' anonymized X-ray images has been built.
8 papers · 2 benchmarks
FeTS2022 (Federated Tumor Segmentation Challenge 2022)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
8 papers · 0 benchmarks
MISAW (MIcro-Surgical Anastomose Workflow recognition on training sessions)
The MISAW data set is composed of 27 sequences of micro-surgical anastomosis on artificial blood vessels performed by 3 surgeons and 3 engineering students.
8 papers · 1 benchmark
MyoPS is a dataset for myocardial pathology segmentation combining three-sequence cardiac magnetic resonance (CMR) images, which was first proposed in the MyoPS challenge, in conjunction with MICCAI 2020.
8 papers · 0 benchmarks
Phee is a dataset for pharmacovigilance comprising over 5000 annotated events from medical case reports and biomedical literature.
8 papers · 0 benchmarks
Over 1.5K images selected from the public Kaggle DR Detection dataset; Five DR grades (DR0 / DR1 / DR2 / DR3 / DR4), re-labeled by a panel of 45 experienced ophthalmologists; Eight retinal lesion classes, including microaneurysm,…
8 papers · 0 benchmarks
Histopathological characterization of colorectal polyps allows to tailor patients' management and follow up with the ultimate aim of avoiding or promptly detecting an invasive carcinoma.
8 papers · 0 benchmarks
Prediction of Finger Flexion IV Brain-Computer Interface Data Competition The goal of this dataset is to predict the flexion of individual fingers from signals recorded from the surface of the brain (electrocorticography (ECoG)).
7 papers · 1 benchmark
The CheXmask Database presents a comprehensive, uniformly annotated collection of chest radiographs, constructed from five public databases: ChestX-ray8, Chexpert, MIMIC-CXR-JPG, Padchest and VinDr-CXR.
7 papers · 0 benchmarks
DFUC2021 (Diabetic Foot Ulcers 2021)
The Diabetic Foot Ulcers dataset (DFUC2021) is a dataset for analysis of pathology, focusing on infection and ischaemia.
7 papers · 0 benchmarks
ISRUC-Sleep is a polysomnographic (PSG) dataset.
7 papers · 2 benchmarks
The KUMC dataset for polyp detection and classification was collected from the University of Kansas Medical Center.
7 papers · 0 benchmarks
KiTS19 (The 2019 Kidney and Kidney Tumor Segmentation Challenge)
The 2021 Kidney and Kidney Tumor Segmentation challenge (abbreviated KiTS21) is a competition in which teams compete to develop the best system for automatic semantic segmentation of renal tumors and surrounding anatomy.
7 papers · 1 benchmark
The Medical Dataset for Abbreviation Disambiguation for Natural Language Understanding (MeDAL) is a large medical text dataset curated for abbreviation disambiguation, designed for natural language understanding pre-training in the medical…
7 papers · 0 benchmarks
The MedDialog dataset (Chinese) contains conversations (in Chinese) between doctors and patients.
7 papers · 0 benchmarks
PHM2017 is a new dataset consisting of 7,192 English tweets across six diseases and conditions: Alzheimer’s Disease, heart attack (any severity), Parkinson’s disease, cancer (any type), Depression (any severity), and Stroke.
7 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.