Home › Datasets › modality › Medical
Medical datasets
archive 2025-07-28
394 datasets carry the modality tag "Medical", ordered by the archive's paper count. Page 6 of 9: 48 shown of 394. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Medical datasets 241–288 of 394
The MIMIC PERform Testing dataset contains the following physiological signals recorded from 200 critically-ill patients during routine clinical care: - electrocardiogram (ECG) - photoplethysmogram (PPG) - impedance pneumography (imp),…
2 papers · 2 benchmarks
Our primary objective in creating this dataset is to support researchers in the advancement of algorithms for keypoints detection and the pretraining of large models on retinal images using a self-supervised approach.
2 papers · 0 benchmarks
MedMNIST-C is an open-source data set collection comprising algorithmically generated corruptions applied to the test sets of the MedMNIST collection following the concept of ImageNet-C.
2 papers · 0 benchmarks
This collection contains images from 422 non-small cell lung cancer (NSCLC) patients.
2 papers · 0 benchmarks
The National Lung Screening Trial (NLST) was a randomized controlled trial conducted by the Lung Screening Study group (LSS) and the American College of Radiology Imaging Network (ACRIN) to determine whether screening for lung cancer with…
2 papers · 1 benchmark
Abstract The Norwegian Endurance Athlete ECG Database contains 12-lead ECG recordings from 28 elite athletes from various sports in Norway.
2 papers · 0 benchmarks
OADAT (OADAT: Experimental and Synthetic Clinical Optoacoustic Data for Standardized Image Processing)
An experimental and synthetic (simulated) OA raw signals and reconstructed image domain datasets rendered with different experimental parameters and tomographic acquisition geometries.
2 papers · 0 benchmarks
PAX-Ray++ (Projected Anatomy in X-Ray Dataset ++)
The PAX-Ray++ dataset uses pseudo-labeled thorax CTs to enable the segmentation of anatomy in Chest X-Rays.
2 papers · 0 benchmarks
Dataset Card for The Cancer Genome Atlas (TCGA) Multimodal Dataset The Cancer Genome Atlas (TCGA) Multimodal Dataset is a comprehensive collection of clinical data, pathology reports, molecular, and slide images for cancer patients.
2 papers · 0 benchmarks
Introduction The 2016 PhysioNet/CinC Challenge aims to encourage the development of algorithms to classify heart sound recordings collected from a variety of clinical or nonclinical (such as in-home visits) environments.
2 papers · 0 benchmarks
PulseImpute is a benchmark for Pulsative Physiological Signal Imputation which includes realistic mHealth missingness models, an extensive set of baselines, and clinically-relevant downstream tasks.
2 papers · 0 benchmarks
RETOUCH (RETOUCH -The Retinal OCT Fluid Detection and Segmentation Benchmark and Challenge)
The goal of the challenge is to compare automated algorithms that are able to detect and segment various types of fluids on a common dataset of optical coherence tomography (OCT) volumes representing different retinal diseases, acquired…
2 papers · 0 benchmarks
RSDD-Time is a dataset of 598 manually annotated self-reported depression diagnosis posts from Reddit that include temporal information about the diagnosis.
2 papers · 0 benchmarks
SemClinBr (A multi‑institutional and multi‑specialty semantically annotated corpus for Portuguese clinical NLP tasks)
Background: The high volume of research focusing on extracting patient information from electronic health records (EHRs) has led to an increase in the demand for annotated corpora, which are a precious resource for both the development and…
2 papers · 1 benchmark
This mouse cerebellar atlas can be used for mouse cerebellar morphometry.
2 papers · 0 benchmarks
The ULS23 test set contains 725 lesions from 284 patients of the Radboudumc and JBZ hospitals in the Netherlands.
2 papers · 1 benchmark
Ward2ICU is a vital signs dataset of inpatients from the general ward.
2 papers · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this DIF file.
1 paper · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this Excel file.
1 paper · 0 benchmarks
ABCD Study (Adolescent Brain Cognitive Development)
The ABCD Study is a prospective longitudinal study starting at the ages of 9-10 and following participants for 10 years.
1 paper · 0 benchmarks
ACCT Data Repository (ACCT is a fast and accessible automatic cell counting tool using machine learning for 2D image segmentation)
This dataset is a collection of fluorescent images from mice in order to test an automatic cell counting tool that we developed.
1 paper · 0 benchmarks
Dataset Card for the ACR Appropriateness Criteria Corpus This dataset contains chunked guidelines and narratives from the ACR Appropriateness Criteria, an set of societal guidelines from the American College of Radiology (ACR) to help…
1 paper · 0 benchmarks
We introduce a new AI-ready computational pathology dataset containing restained and co-registered digitized images from eight head-and-neck squamous cell carcinoma patients.
1 paper · 0 benchmarks
AIROGS (Rotterdam EyePACS AIROGS)
The Rotterdam EyePACS AIROGS dataset (in full, so including train and test) contains 113,893 color fundus images from 60,357 subjects and approximately 500 different sites with a heterogeneous ethnicity.
1 paper · 0 benchmarks
The AneuX morphology database includes data from 3 different data sources: AneuX, @neurIST and Aneurisk.
1 paper · 0 benchmarks
BCSS (Breast Cancer Semantic Segmentation)
The BCSS dataset contains over 20,000 segmentation annotations of tissue regions from breast cancer images from The Cancer Genome Atlas (TCGA).
1 paper · 0 benchmarks
This dataset is a BIDS-compatible version of the CHB-MIT Scalp EEG Database.
1 paper · 0 benchmarks
This dataset is a BIDS compatible version of the Siena Scalp EEG Database.
1 paper · 0 benchmarks
Our dataset, BSMDD, was collected from various open social media platforms and translated and annotated by native Bengali speakers with expertise in both language and mental health.
1 paper · 0 benchmarks
Overview This is a dataset of blood cells photos.
1 paper · 0 benchmarks
BraTS PEDs 2023 (The Brain Tumor Segmentation (BraTS) Challenge 2023: Focus on Pediatrics (CBTN-CONNECT-DIPGR-ASNR-MICCAI BraTS-PEDs))
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
BreastDICOM4 ([MIMBCD-UI] UTA4: Medical Imaging DICOM Files Dataset)
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 1 benchmark
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 0 benchmarks
The CLOUD dataset is a set of Optical Coherence Tomography of the Anterior Segment images (AS-OCT) used to the automatic identification and representation of the cornea-contact lens relationship.
1 paper · 0 benchmarks
CMeIE (Chinese Medical Information Extraction Dataset)
Chinese Medical Information Extraction, a dataset that is also released in CHIP2020, is used for CMeIE task.
1 paper · 1 benchmark
COVIDx CXR-3 is an open access benchmark dataset that we generated, comprising 30,882 CXR images across 17,026 patient cases.
1 paper · 1 benchmark
CPCXR (COVID-19 Posteroanterior Chest X-Ray fused)
The COVID-19 Posteroanterior Chest X-Ray fused (CPCXR) dataset is generated by the fusion of three publicly available datasets: COVID-19 cxr image, Radiological Society of North America (RSNA), and U.S.
1 paper · 0 benchmarks
CPSC2019 (The 2nd China Physiological Signal Challenge (CPSC 2019))
Introduction The China Physiological Signal Challenge 2019 (CPSC 2019) aims to encourage the development of algorithms for challenging QRS detection and heart rate (HR) estimation from short-term single-lead ECG recordings usually with low…
1 paper · 0 benchmarks
CPSC2020 (The 3rd China Physiological Signal Challenge 2020)
Introduction Abnormality of cardiac conduction system can induce arrhythmia.
1 paper · 0 benchmarks
CPSC2021 (The 4th China Physiological Signal Challenge 2021)
Introduction The 4th China Physiological Signal Challenge 2021 (CPSC 2021) aims to encourage the development of algorithms for searching the paroxysmal atrial fibrillation (PAF) events from dynamic ECG recordings.
1 paper · 0 benchmarks
Histological images of colorectal cancer, derived from the TCGA database
1 paper · 0 benchmarks
The dataset has 93 image stacks and their corresponding Extended Depth of Field (EDF) image acquired from cases with grades Nagative, LSIL or HSIL (The Bethesda System): - Negative: 16 - LSIL: 46 - HSIL: 31 The ground truth includes the…
1 paper · 0 benchmarks
CoCaHis (Colon Cancer Histology Dataset)
Highlights • Publicly available dataset with 82 H&E stained images of frozen sections.
1 paper · 0 benchmarks
The data set includes 589 T2-weighted images acquired from the same number of patients collected by seven studies, INDEX, the SmartTarget Biopsy Trial, PICTURE, TCIA Prostate3T, Promise12, TCIA ProstateDx (Diagnosis) and the Prostate MR…
1 paper · 0 benchmarks
Data Set Information: The main goal of this data set is providing clean and valid signals for designing cuff-less blood pressure estimation algorithms.
1 paper · 0 benchmarks
DLBCL-Morph is a dataset containing 42 digitally scanned high-resolution tissue microarray (TMA) slides accompanied by clinical, cytogenetic, and geometric features from 209 DLBCL cases.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.