Home › Datasets › modality › Biomedical

Biomedical datasets

archive 2025-07-28

122 datasets carry the modality tag "Biomedical", ordered by the archive's paper count. Page 1 of 3: 48 shown of 122. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Biomedical datasets 1–48 of 122

Kvasir-SEG is an open-access dataset of gastrointestinal polyp images and corresponding segmentation masks, manually annotated by a medical doctor and then verified by an experienced gastroenterologist.
201 papers · 2 benchmarks
BLUE (Biomedical Language Understanding Evaluation)
The BLUE benchmark consists of five different biomedicine text-mining tasks with ten corpora.
133 papers · 0 benchmarks
ACDC (Automated Cardiac Diagnosis Challenge)
The goal of the Automated Cardiac Diagnosis Challenge (ACDC) challenge is to: - compare the performance of automatic methods on the segmentation of the left ventricular endocardium and epicardium as the right ventricular endocardium for…
52 papers · 5 benchmarks
Dataset contains 33,010 molecule-description pairs split into 80\%/10\%/10\% train/val/test splits.
43 papers · 4 benchmarks
BLURB (Biomedical Language Understanding and Reasoning Benchmark)
BLURB is a collection of resources for biomedical natural language processing.
40 papers · 2 benchmarks
Contains hundreds of frontal view X-rays and is the largest public resource for COVID-19 image and prognostic data, making it a necessary resource to develop and evaluate tools to aid in the treatment of COVID-19.
35 papers · 1 benchmark
BioGRID (Biological General Repository for Interaction Datasets)
BioGRID is a biomedical interaction repository with data compiled through comprehensive curation efforts.
34 papers · 2 benchmarks
MS-CXR (Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing)
The MS-CXR dataset provides 1162 image–sentence pairs of bounding boxes and corresponding phrases, collected across eight different cardiopulmonary radiological findings, with an approximately equal number of pairs for each finding.
32 papers · 0 benchmarks
ATOM3D is a unified collection of datasets concerning the three-dimensional structure of biomolecules, including proteins, small molecules, and nucleic acids.
23 papers · 0 benchmarks
ICBHI Respiratory Sound Database (The Respiratory Sound database - ICBHI 2017 Challenge)
The Respiratory Sound database was originally compiled to support the scientific challenge organized at Int.
21 papers · 2 benchmarks
SICAPv2 is a database containing prostate histology whole slide images with both annotations of global Gleason scores and path-level Gleason grades.
21 papers · 0 benchmarks
BCI (Breast Cancer Immunohistochemical Image Generation)
The evaluation of human epidermal growth factor receptor 2 (HER2) expression is essential to formulate a precise treatment for breast cancer.
19 papers · 1 benchmark
The National Institutes of Health’s Clinical Center has made a large-scale dataset of CT images publicly available to help the scientific community improve detection accuracy of lesions.
19 papers · 1 benchmark
Consists of annotated frames containing GI procedure tools such as snares, balloons and biopsy forceps, etc.
19 papers · 3 benchmarks
LIVECell (Label-free In Vitro image Examples of Cells)
The LIVECell (Label-free In Vitro image Examples of Cells) dataset is a large-scale microscopic image dataset for instance-segmentation of individual cells in 2D cell cultures.
18 papers · 1 benchmark
ChemProt consists of 1,820 PubMed abstracts with chemical-protein interactions annotated by domain experts and was used in the BioCreative VI text mining chemical-protein interactions shared task.
16 papers · 1 benchmark
IDRiD (Indian Diabetic Retinopathy Image Dataset)
Indian Diabetic Retinopathy Image Dataset (IDRiD) dataset consists of typical diabetic retinopathy lesions and normal retinal structures annotated at a pixel level.
16 papers · 3 benchmarks
BioLAMA is a benchmark comprised of 49K biomedical factual knowledge triples for probing biomedical Language Models.
14 papers · 0 benchmarks
HyperKvasir dataset contains 110,079 images and 374 videos where it captures anatomical landmarks and pathological and normal findings.
12 papers · 2 benchmarks
Kaggle EyePACS (Kaggle EyePACS. Diabetic Retinopathy Detection Identify signs of diabetic retinopathy in eye images)
Diabetic retinopathy is the leading cause of blindness in the working-age population of the developed world.
11 papers · 1 benchmark
UI-PRMD (University of Idaho – Physical Rehabilitation Movement Dataset)
UI-PRMD is a data set of movements related to common exercises performed by patients in physical therapy and rehabilitation programs.
10 papers · 2 benchmarks
CHAOS (CHAOS - Combined (CT-MR) Healthy Abdominal Organ Segmentation)
CHAOS challenge aims the segmentation of abdominal organs (liver, kidneys and spleen) from CT and MRI data.
9 papers · 0 benchmarks
These are 10 synthetic genomics datasets generated with NEAT v3 (based on TP53 gene of Homo Sapiens) for the use case of benchmarking somatic variant callers.
8 papers · 1 benchmark
ATLAS v2.0 (Anatomical Tracings of Lesions After Stroke Dataset version 2.0)
Accurate lesion segmentation is critical in stroke rehabilitation research for the quantification of lesion burden and accurate image processing.
8 papers · 1 benchmark
NuCLS (Nucleus Classification, Localization and Segmentation)
The NuCLS dataset contains over 220,000 labeled nuclei from breast cancer images from TCGA.
8 papers · 0 benchmarks
Phee is a dataset for pharmacovigilance comprising over 5000 annotated events from medical case reports and biomedical literature.
8 papers · 0 benchmarks
Prediction of Finger Flexion IV Brain-Computer Interface Data Competition The goal of this dataset is to predict the flexion of individual fingers from signals recorded from the surface of the brain (electrocorticography (ECoG)).
7 papers · 1 benchmark
The CheXmask Database presents a comprehensive, uniformly annotated collection of chest radiographs, constructed from five public databases: ChestX-ray8, Chexpert, MIMIC-CXR-JPG, Padchest and VinDr-CXR.
7 papers · 0 benchmarks
Language-molecule models have emerged as an exciting direction for molecular discovery and understanding.
7 papers · 1 benchmark
MIPE (Improving Paratope and Epitope Prediction by Multi-Modal Contrastive Learning and Interaction Informativeness Estimation)
Datasets.
7 papers · 1 benchmark
EHR-RelB is a benchmark dataset for biomedical concept relatedness, consisting of 3630 concept pairs sampled from electronic health records (EHRs).
6 papers · 0 benchmarks
FrenchMedMCQA (FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain)
This paper introduces FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain.
6 papers · 1 benchmark
PLABA (Plain Language Adaptation of Biomedical Abstracts)
Plain Language Adaptation of Biomedical Abstracts (PLABA) is a dataset designed for automatic adaptation that is both document- and sentence-aligned.
6 papers · 0 benchmarks
PhysioNet Challenge 2021 (The PhysioNet/Computing in Cardiology Challenge 2021)
Data Description The training data contains twelve-lead ECGs.
6 papers · 2 benchmarks
The 2017 PhysioNet/CinC Challenge aims to encourage the development of algorithms to classify, from a single short ECG lead recording (between 30 s and 60 s in length), whether the recording shows normal sinus rhythm, atrial fibrillation…
5 papers · 0 benchmarks
CBC (Complete Blood Count)
The complete blood count (CBC) dataset contains 360 blood smear images along with their annotation files splitting into Training, Testing, and Validation sets.
5 papers · 0 benchmarks
Benchmark for de novo molecular design
5 papers · 0 benchmarks
Kvasir-Sessile dataset (Sessile polyps from Kvasir-SEG)
The Kvasir-SEG dataset includes 196 polyps smaller than 10 mm classified as Paris class 1 sessile or Paris class IIa.
5 papers · 0 benchmarks
LUDB (Lobachevsky University Electrocardiography Database)
Abstract Lobachevsky University Electrocardiography Database (LUDB) is an ECG signal database with marked boundaries and peaks of P, T waves and QRS complexes.
5 papers · 1 benchmark
This data collection consists of images acquired during chemoradiotherapy of 20 locally-advanced, non-small cell lung cancer patients.
5 papers · 0 benchmarks
MICCAI Challenge on Circuit Reconstruction from Electron Microscopy Images.
4 papers · 1 benchmark
CoVERT (A Corpus of Fact-checked Biomedical COVID-19 Tweets)
CoVERT is a fact-checked corpus of tweets with a focus on the domain of biomedicine and COVID-19-related (mis)information.
4 papers · 0 benchmarks
LIMUC (Labeled Images for Ulcerative Colitis)
The LIMUC dataset is the largest publicly available labeled ulcerative colitis dataset that compromises 11276 images from 564 patients and 1043 colonoscopy procedures.
4 papers · 1 benchmark
Data The data for this Challenge are from multiple sources: CPSC Database and CPSC-Extra Database INCART Database PTB and PTB-XL Database The Georgia 12-lead ECG Challenge (G12EC) Database Undisclosed Database The first source is the…
4 papers · 1 benchmark
The eSports Sensors dataset contains sensor data collected from 10 players in 22 matches in League of Legends.
4 papers · 2 benchmarks
DIPS-Plus (The Enhanced Database of Interacting Protein Structures for Interface Prediction)
How and where proteins interface with one another can ultimately impact the proteins' functions along with a range of other biological processes.
3 papers · 0 benchmarks
The dataset contains a Video capsule endoscopy dataset for polyp segmentation.
3 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.