Home › Datasets › modality › Medical
Medical datasets
archive 2025-07-28
394 datasets carry the modality tag "Medical", ordered by the archive's paper count. Page 2 of 9: 48 shown of 394. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Medical datasets 49–96 of 394
MeQSum is a dataset for medical question summarization.
33 papers · 1 benchmark
CholecT50 is a dataset of endoscopic videos of laparoscopic cholecystectomy surgery introduced to enable research on fine-grained action recognition in laparoscopic surgery.
32 papers · 5 benchmarks
MS-CXR (Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing)
The MS-CXR dataset provides 1162 image–sentence pairs of bounding boxes and corresponding phrases, collected across eight different cardiopulmonary radiological findings, with an approximately equal number of pairs for each finding.
32 papers · 0 benchmarks
MedQuAD (Medical Question Answering Dataset)
MedQuAD includes 47,457 medical question-answer pairs created from 12 NIH websites (e.g.
32 papers · 0 benchmarks
LC25000 (Lung And Colon Histopathological Image Dataset)
The LC25000 dataset contains 25,000 color images with 5 classes of 5,000 images each.
31 papers · 0 benchmarks
Under Institutional Review Board (IRB) supervision, 50 abdomen CT scans of were randomly selected from a combination of an ongoing colorectal cancer chemotherapy trial, and a retrospective ventral hernia study.
31 papers · 3 benchmarks
The MIT-BIH Arrhythmia Database contains 48 half-hour excerpts of two-channel ambulatory ECG recordings, obtained from 47 subjects studied by the BIH Arrhythmia Laboratory between 1975 and 1979.
31 papers · 5 benchmarks
BRACS (BReAst Carcinoma Subtyping)
BReAst Carcinoma Subtyping (BRACS) dataset, a large cohort of annotated Hematoxylin & Eosin (H&E)-stained images to facilitate the characterization of breast lesions.
30 papers · 0 benchmarks
MedICaT is a dataset of medical images, captions, subfigure-subcaption annotations, and inline textual references.
28 papers · 0 benchmarks
SegTHOR (Segmentation of THoracic Organs at Risk)
SegTHOR (Segmentation of THoracic Organs at Risk) is a dataset dedicated to the segmentation of organs at risk (OARs) in the thorax, i.e.
28 papers · 0 benchmarks
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
27 papers · 1 benchmark
Head and Neck Tumor Segmentation
26 papers · 0 benchmarks
ROSE (Retinal OCTA SEgmentation dataset)
Retinal OCTA SEgmentation dataset (ROSE) consists of 229 OCTA images with vessel annotations at either centerline-level or pixel level.
25 papers · 4 benchmarks
MMSE-HR (Multimodal Spontaneous Expression-Heart Rate dataset)
The MMSE-HR benchmark consists of a dataset of 102 videos from 40 subjects recorded at 1040x1392 raw resolution at 25fps.
24 papers · 1 benchmark
IXI (IXI Brain Development Dataset)
IXI Dataset is a collection of 600 MR brain images from normal, healthy subjects.
23 papers · 4 benchmarks
The PhysioNet Challenge 2012 dataset is publicly available and contains the de-identified records of 8000 patients in Intensive Care Units (ICU).
23 papers · 5 benchmarks
TCGA (The Cancer Genome Atlas)
23 papers · 2 benchmarks
Retrospectively collected medical data has the opportunity to improve patient care through knowledge discovery and algorithm development.
22 papers · 0 benchmarks
MosMedData contains anonymised human lung computed tomography (CT) scans with COVID-19 related findings, as well as without such findings.
22 papers · 1 benchmark
The ECGs in this collection were obtained using a non-commercial, PTB prototype recorder with the following specifications: 16 input channels, (14 for ECGs, 1 for respiration, 1 for line voltage) Input voltage: ±16 mV, compensated offset…
22 papers · 4 benchmarks
CliCR is a new dataset for domain specific reading comprehension used to construct around 100,000 cloze queries from clinical case reports.
21 papers · 1 benchmark
The Respiratory Sound database was originally compiled to support the scientific challenge organized at Int.
21 papers · 2 benchmarks
SICAPv2 is a database containing prostate histology whole slide images with both annotations of global Gleason scores and path-level Gleason grades.
21 papers · 0 benchmarks
This dataset has 1,842 images with pixel-level DR-related lesion annotations, and 1,000 images with image-level labels graded by six board-certified ophthalmologists with intra-rater consistency.
20 papers · 0 benchmarks
BCI (Breast Cancer Immunohistochemical Image Generation)
The evaluation of human epidermal growth factor receptor 2 (HER2) expression is essential to formulate a precise treatment for breast cancer.
19 papers · 1 benchmark
Chaoyang dataset contains 1111 normal, 842 serrated, 1404 adenocarcinoma, 664 adenoma, and 705 normal, 321 serrated, 840 adenocarcinoma, 273 adenoma samples for training and testing, respectively.
19 papers · 2 benchmarks
DDXPlus (DDXPlus: A New Dataset For Automatic Medical Diagnosis)
There has been a rapidly growing interest in Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the machine learning research literature, aiming to assist doctors in telemedicine services.
19 papers · 0 benchmarks
The National Institutes of Health’s Clinical Center has made a large-scale dataset of CT images publicly available to help the scientific community improve detection accuracy of lesions.
19 papers · 1 benchmark
Consists of annotated frames containing GI procedure tools such as snares, balloons and biopsy forceps, etc.
19 papers · 3 benchmarks
CrossMoDA is a large and multi-class benchmark for unsupervised cross-modality Domain Adaptation.
18 papers · 0 benchmarks
The Endomapper dataset is the first collection of complete endoscopy sequences acquired during regular medical practice, including slow and careful screening explorations, making secondary use of medical data.
18 papers · 0 benchmarks
The dataset for this challenge was obtained by carefully annotating tissue images of several patients with tumors of different organs and who were diagnosed at multiple hospitals.
17 papers · 2 benchmarks
Recent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online.
17 papers · 0 benchmarks
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
IDRiD (Indian Diabetic Retinopathy Image Dataset)
Indian Diabetic Retinopathy Image Dataset (IDRiD) dataset consists of typical diabetic retinopathy lesions and normal retinal structures annotated at a pixel level.
16 papers · 3 benchmarks
SKM-TEA (Stanford Knee MRI with Multi-Task Evaluation)
The SKM-TEA dataset pairs raw quantitative knee MRI (qMRI) data, image data, and dense labels of tissues and pathology for end-to-end exploration and evaluation of the MR imaging pipeline.
16 papers · 0 benchmarks
The ISIC 2017 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
15 papers · 0 benchmarks
The MSK dataset is a dataset for lesion recognition from the Memorial Sloan-Kettering Cancer Center.
15 papers · 0 benchmarks
MedVidQA (Medical Video Question Answering)
The MedVidQA dataset contains the collection of 3, 010 manually created health-related questions and timestamps as visual answers to those questions from trusted video sources, such as accredited medical schools with an established…
15 papers · 0 benchmarks
MMPD (Multi-Domain Mobile Video Physiology Dataset)
The Multi-domain Mobile Video Physiology Dataset (MMPD), comprising 11 hours(1152K frames) of recordings from mobile phones of 33 subjects.
14 papers · 0 benchmarks
REFUGE Challenge provides a data set of 1200 fundus images with ground truth segmentations and clinical glaucoma labels, currently the largest existing one.
14 papers · 4 benchmarks
Purpose Medical imaging has become increasingly important in diagnosing and treating oncological patients, particularly in radiotherapy.
14 papers · 0 benchmarks
A large-scale cloze-style biomedical MRC dataset.
13 papers · 1 benchmark
BRATS 2016 is a brain tumor segmentation dataset.
13 papers · 0 benchmarks
BrixIA Covid-19 is a large dataset of CXR images corresponding to the entire amount of images taken for both triage and patient monitoring in sub-intensive and intensive care units during one month (between March 4th and April 4th 2020) of…
13 papers · 0 benchmarks
CBIS-DDSM (Curated Breast Imaging Subset of Digital Database for Screening Mammography)
This CBIS-DDSM (Curated Breast Imaging Subset of DDSM) is an updated and standardized version of the Digital Database for Screening Mammography (DDSM) .
13 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.