Home › Datasets › modality › Medical

Medical datasets

archive 2025-07-28

394 datasets carry the modality tag "Medical", ordered by the archive's paper count. Page 4 of 9: 48 shown of 394. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Medical datasets 145–192 of 394

Electrocardiography (ECG) is a key diagnostic tool to assess the cardiac condition of a patient.
7 papers · 2 benchmarks
English subset of the SLAKE dataset, comprising 642 images and more than 7,000 question–answer pairs.
7 papers · 0 benchmarks
The m2cai16-tool-locations dataset contains spatial tool annotations for 2,532 frames across the first 10 videos in the m2cai16-tool dataset, which includes 15 videos in total.
7 papers · 0 benchmarks
The 3DSeg-8 is a collection of several publicly available 3D segmentation datasets from different medical imaging modalities, e.g.
6 papers · 0 benchmarks
ARCADE (Automatic Region-based Coronary Artery Disease diagnostics using x-ray angiography imagEs Dataset)
ARCADE: Automatic Region-based Coronary Artery Disease diagnostics using x-ray angiography imagEs Dataset Phase 2 consist of two folders with 300 images in each of them as well as annotations.
6 papers · 0 benchmarks
CHB-MIT (CHB-MIT Scalp EEG)
The CHB-MIT dataset is a dataset of EEG recordings from pediatric subjects with intractable seizures.
6 papers · 1 benchmark
A dataset of 12-lead ECGs with annotations.
6 papers · 1 benchmark
DRTiD is a benchmark dataset for DR grading, consisting of 3,100 two-field fundus images.
6 papers · 0 benchmarks
FrenchMedMCQA (FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain)
This paper introduces FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain.
6 papers · 1 benchmark
GBCU (Gallbladder Cancer Ultrasound Dataset)
GBCU is the first public dataset for Gallbladder Cancer identification from Ultrasound images.
6 papers · 1 benchmark
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
6 papers · 0 benchmarks
The RAD-ChestCT dataset is a large medical imaging dataset developed by Duke MD/PhD Rachel Draelos during her Computer Science PhD supervised by Lawrence Carin.
6 papers · 0 benchmarks
BRATS 2014 is a brain tumor segmentation dataset.
5 papers · 1 benchmark
This brain anatomy segmentation dataset has 1300 2D US scans for training and 329 for testing.
5 papers · 1 benchmark
CBC (Complete Blood Count)
The complete blood count (CBC) dataset contains 360 blood smear images along with their annotation files splitting into Training, Testing, and Validation sets.
5 papers · 0 benchmarks
Request access: cadpath.ai@impdiagnostics.com The CRC dataset contains 1133 colorectal biopsy and polypectomy slides and is the result of our ongoing efforts to contribute to CRC diagnosis with a reference dataset.
5 papers · 0 benchmarks
Cata7 is the first cataract surgical instrument dataset for semantic segmentation.
5 papers · 0 benchmarks
HaN-Seg (The Head and Neck Organ-at-Risk CT & MR Segmentation Challenge)
Cancer in the region of the head and neck (HaN) is one of the most prominent cancers, for which radiotherapy represents an important treatment modality that aims to deliver a high radiation dose to the targeted cancerous cells while…
5 papers · 0 benchmarks
The ISIC 2017 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
5 papers · 0 benchmarks
Kvasir-Sessile dataset (Sessile polyps from Kvasir-SEG)
The Kvasir-SEG dataset includes 196 polyps smaller than 10 mm classified as Paris class 1 sessile or Paris class IIa.
5 papers · 0 benchmarks
Lesion Boundary Segmentation Dataset is a dataset for lesion segmentation from the ISIC2018 challenge.
5 papers · 0 benchmarks
MESSIDOR (MESSIDOR DATABASE)
The Messidor database has been established to facilitate studies on computer-assisted diagnoses of diabetic retinopathy.
5 papers · 0 benchmarks
NIH-CXR-LT (Long-tailed (LT) NIH ChestXRay14)
NIH-CXR-LT.
5 papers · 1 benchmark
RITE (Retinal Images vessel Tree Extraction)
The RITE (Retinal Images vessel Tree Extraction) is a database that enables comparative studies on segmentation or classification of arteries and veins on retinal fundus images, which is established based on the public available DRIVE…
5 papers · 2 benchmarks
The Raider dataset collects fMRI recordings of 1000 voxels from the ventral temporal cortex, for 10 healthy adult participants passively watching the full-length movie “Raiders of the Lost Ark”.
5 papers · 0 benchmarks
SERV-CT (SERV-CT: A disparity dataset from CT for validation of endoscopic 3D reconstruction)
Endoscopic stereo reconstruction for surgical scenes gives rise to specific problems, including the lack of clear corner features, highly specular surface properties, and the presence of blood and smoke.
5 papers · 0 benchmarks
Thyroid (Thyroid Disease)
Thyroid is a dataset for detection of thyroid diseases, in which patients diagnosed with hypothyroid or subnormal are anomalies against normal patients.
5 papers · 1 benchmark
VinDr-RibCXR is a benchmark dataset for automatic segmentation and labeling of individual ribs from chest X-ray (CXR) scans.
5 papers · 0 benchmarks
A whole-body FDG-PET/CT dataset with manually annotated tumor lesions (FDG-PET-CT-Lesions) 1,014 studies (900 patients)
4 papers · 0 benchmarks
CODA-19 is a human-annotated dataset that denotes the Background, Purpose, Method, Finding/Contribution, and Other for 10,966 English abstracts in the COVID-19 Open Research Dataset.
4 papers · 0 benchmarks
DiagSet is a histopathological dataset for prostate cancer detection.
4 papers · 0 benchmarks
ESAD (SARAS Endoscopic Surgeon Action Detection)
ESAD is a large-scale dataset designed to tackle the problem of surgeon action detection in endoscopic minimally invasive surgery.
4 papers · 0 benchmarks
FIRE (Fundus Image Registration Dataset)
Fundus Image Registration Dataset (FIRE) is a dataset consisting of 129 retinal images forming 134 image pairs.
4 papers · 1 benchmark
The IS-A dataset is a dataset of relations extracted from a medical ontology.
4 papers · 0 benchmarks
LIMUC (Labeled Images for Ulcerative Colitis)
The LIMUC dataset is the largest publicly available labeled ulcerative colitis dataset that compromises 11276 images from 564 patients and 1043 colonoscopy procedures.
4 papers · 1 benchmark
MIMIC-CXR-LT (long-tailed version of MIMIC-CXR)
MIMIC-CXR-LT.
4 papers · 1 benchmark
MIMIC-IV-ED is a large, freely available database of emergency department (ED) admissions at the Beth Israel Deaconess Medical Center between 2011 and 2019.
4 papers · 0 benchmarks
For each dataset we provide a short description as well as some characterization metrics.
4 papers · 0 benchmarks
OVQA contains 19,020 medical visual question and answer pairs generated from 2,001 medical images collected from 2,212 EMRs in Orthopedics.
4 papers · 0 benchmarks
QT-NSTDB (QT database + MIT-BIH Noise Stress Test Database (NSTDB))
We designed a baseline wander (BLW) removal benchmark to evaluate various methods using a consistent test set and uniform conditions.
4 papers · 1 benchmark
VietMed (VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain)
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
4 papers · 2 benchmarks
eICU-CRD (eICU Collaborative Research Database)
The eICU Collaborative Research Database is a large multi-center critical care database made available by Philips Healthcare in partnership with the MIT Laboratory for Computational Physiology.
4 papers · 2 benchmarks
Attention Deficit Hyperactivity Disorder (ADHD) affects at least 5-10% of school-age children and is associated with substantial lifelong impairment, with annual direct costs exceeding $36 billion/year in the US.
3 papers · 0 benchmarks
ATM'22 is a multi-site, multi-domain dataset for pulmonary airway segmentation.
3 papers · 0 benchmarks
AeroPath (AeroPath: An airway segmentation benchmark dataset with challenging pathology)
Public benchmark dataset (AeroPath), consisting of 27 CT images from patients with pathologies ranging from emphysema to large tumors, with corresponding trachea and bronchi annotations.
3 papers · 0 benchmarks
Apnea-ECG (PhysioNet Apnea-ECG Database)
The data consist of 70 records, divided into a learning set of 35 records (a01 through a20, b01 through b05, and c01 through c10), and a test set of 35 records (x01 through x35), all of which may be downloaded from this page.
3 papers · 1 benchmark
The AxonEM dataset consists of two 30x30x30 um^3 EM image volumes from the human and mouse cortex, respectively.
3 papers · 0 benchmarks
✔️Abstract A Brain tumor is considered as one of the aggressive diseases, among children and adults.
3 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.