Home › Datasets › modality › Medical
Medical datasets
archive 2025-07-28
394 datasets carry the modality tag "Medical", ordered by the archive's paper count. Page 7 of 9: 48 shown of 394. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Medical datasets 289–336 of 394
A dataset of 100K synthetic images of skin lesions, ground-truth (GT) segmentations of lesions and healthy skin, GT segmentations of seven body parts (head, torso, hips, legs, feet, arms and hands), and GT binary masks of non-skin regions…
1 paper · 0 benchmarks
EBHI-Seg is a dataset containing 5,170 images of six types of tumor differentiation stages and the corresponding ground truth images.
1 paper · 0 benchmarks
EGC-FPHFS (Early Gastric Cancer Data from First People's Hospital of Foshan)
High-resolution early gastric cancer (EGC) detection and analysis: Patient Data:Datasets often include images from patients diagnosed with gastric cancer, specifically distinguishing between early gastric cancer (EGC) and Non -pathogenic…
1 paper · 1 benchmark
ENSeg Dataset Overview This dataset represents an enhanced subset of the ENS dataset.
1 paper · 1 benchmark
This is the supplemental data for our paper on how to benchmark registrations of serial sections with ground truths.
1 paper · 0 benchmarks
The dataset X of this work is an extension of the heartSeg dataset.
1 paper · 1 benchmark
Facial Skeletal angles (Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis))
Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis)
1 paper · 0 benchmarks
The Fraunhofer Portugal AICOS EDoF Dataset was produced within the TAMI project and is composed of images of microscopic fields of view (FOV) of Liquid-based Cervical Cytology (LBC) samples.
1 paper · 0 benchmarks
GOD (Generic Object Decoding)
The Generic Object Decoding (GOD) Dataset is a specialized resource developed for fMRI-based decoding.
1 paper · 1 benchmark
Human fibrosarcoma HT1080WT (ATCC) cells at low cell densities embedded in 3D collagen type I matrices [1].
1 paper · 0 benchmarks
HTDM (Hypertention Disease Medication)
Hypertention Disease Medication dataset.
1 paper · 0 benchmarks
HYPE (PPG and Blood Pressure from a Hypertensive Population)
HYPE Dataset - Version 1.0.0 REFERENCE PAPER ------------------- Morassi Sasso, A., Datta, S., Jeitler, M., Steckhan, N., Kessler, C.
1 paper · 0 benchmarks
HuSHeM (Human Sperm Head Morphology Dataset)
At the Isfahan Fertility and Infertility Center, semen samples were collected from fifteen patients.
1 paper · 0 benchmarks
A dataset of A 3D Computed Tomography (CT) image dataset, ImageTBAD, for segmentation of Type-B Aortic Dissection is published.
1 paper · 0 benchmarks
We processed 241 pairs of CXR and DES soft tissue images from the JSRT dataset by performing operations like inversion and contrast adjustment to convert these images into negative formats more frequently used in clinical settings.
1 paper · 0 benchmarks
LIRCAD (Inria Liver vessels subbranch anotomical nomenclature labels - "LIRCAD")
The structure for the dataset is as follows : 3DLiverVasculatureProject/ ├── CT/ │ Contains the CT scans ├── Labels/ │ Contains the dual labels for the vessel tree annotations 0 background, 1 Portal vein, 2 hepatic vein ├──…
1 paper · 0 benchmarks
LLM Health Benchmarks Dataset The Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties.
1 paper · 0 benchmarks
LLaVA-Rad MIMIC-CXR features more accurate section extractions from MIMIC-CXR free-text radiology reports.
1 paper · 0 benchmarks
LPBA40 (LONI Probabilistic Brain Atlas)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset consists of more than 16,000 retinal OCT B-scans from 441 cases (Normal: 120, Drusen: 160, CNV: 161) and is acquired at Noor Eye Hospital, Tehran, Iran.
1 paper · 0 benchmarks
This dataset includes sharp-blur pairs of Leishmania image, which is a protozoan parasite microscopy image dataset of Leishmania, obtained from the preserved slides stained with Giemsa.
1 paper · 0 benchmarks
Liver-US (Liver Ultrasound Dataset for Medical Image Classification)
The Liver-US dataset is a comprehensive collection of high-quality ultrasound images of the liver, including both normal and abnormal cases.
1 paper · 1 benchmark
This dataset contains pre-processed versions of datasets introduced in prior works.
1 paper · 0 benchmarks
The data generated from this study are grouped into 3 main types: (1) participant demographic and clinical data, (2) sensor data from the different devices, as well as clinical scores and metadata related to the tasks performed, and (3)…
1 paper · 0 benchmarks
MODA dataset (Massive Online Data Annotation Spindle Dataset)
MODA is a large open-source dataset of high quality, human-scored sleep spindles (5342 spindles, from 180 subjects) that was produced by the Massive Online Data Annotation project.
1 paper · 1 benchmark
1、 Competition name: The 2nd China Society of Image and Graphics (CSIG) Image and Graphics Technology Challenge: MRSpineSeg Challenge: Automated Multi-class Segmentation of Spinal Structures on Volumetric MR Images.
1 paper · 0 benchmarks
MTNeuro is a multi-task neuroimaging benchmark built on volumetric, micrometer-resolution X-ray microtomography images spanning a large thalamocortical section of mouse brain, encompassing multiple cortical and subcortical regions.
1 paper · 0 benchmarks
A new in-context visual question answering dataset encompassing interleaved image and EHR data derived from MIMIC-IV and MIMIC-CXR-JPG databases.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The MedVidCL dataset contains a collection of 6, 617 videos annotated into ‘medical instructional’, ‘medical non-instructional' and ‘non-medical’ classes.
1 paper · 0 benchmarks
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
MediConfusion is a challenging medical Visual Question Answering (VQA) benchmark dataset, that probes the failure modes of medical Multimodal Large Language Models (MLLMs) from a vision perspective.
1 paper · 0 benchmarks
Medical Case Report Corpus is a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library.
1 paper · 0 benchmarks
MuCeD, a dataset that is carefully curated and validated by expert pathologists from the All India Institute of Medical Science (AIIMS), Delhi, India.
1 paper · 0 benchmarks
Mouse Brain MRI atlas (both in-vivo and ex-vivo) (repository relocated from the original webpage) List of atlases - FVBNCrl: Brain MRI atlas of the wild-type FVBNCrl mouse strain (used as the background strain for the rTg4510 which is a…
1 paper · 0 benchmarks
NCANDA (National Consortium on Alcohol and Neurodevelopment in Adolescence)
The NCANDA consortium is composed of an Administrative component at the University of California San Diego, a Data Analysis and Informatics component at SRI International, and five research sites (University of California San Diego, SRI…
1 paper · 0 benchmarks
NIH-Lymph Node (NIH-LN) contains 388 mediastinal LNs in 90 CT scans and 595 abdominal LNs in 86 scans.
1 paper · 0 benchmarks
Unique radiogenomic dataset from a Non-Small Cell Lung Cancer (NSCLC) cohort of 211 subjects.
1 paper · 0 benchmarks
The NVALT-11 study considered the effect of profylactic brain radiation versus observation in (m=174) patients with advanced non-small cell lung cancer.
1 paper · 0 benchmarks
Te NVALT-8 study (m=200 participants) examined if nadroparin combined with chemotherapy could reduce cancer relapse after surgical removal of a non-small cell lung tumour.
1 paper · 0 benchmarks
Onchocerciasis is causing blindness in over half a million people in the world today.
1 paper · 0 benchmarks
Authors of the Dataset: - Pratik Bhowal (B.E., Dept of Electronics and Instrumentation Engineering, Jadavpur University Kolkata, India)…
1 paper · 1 benchmark
The PART-OF dataset is a dataset of relations extracted from a medical ontology.
1 paper · 0 benchmarks
Overview PASSION derm is a pioneering initiative dedicated to closing the diversity gap in dermatology datasets.
1 paper · 0 benchmarks
PRECOG (PREdiction of Clinical Outcomes from Genomic Profiles)
The PREdiction of Clinical Outcomes from Genomic profiles (or PRECOG) encompasses 166 cancer expression data sets, including overall survival data for ~18,000 patients diagnosed with 39 distinct malignancies.
1 paper · 0 benchmarks
PWISeg (PWISeg Surgical Instruments Dataset)
Overview The Surgical Instruments Recognition Dataset is a groundbreaking collection of high-resolution images (1280x960 pixels) specifically designed for the recognition and categorization of surgical instruments.
1 paper · 0 benchmarks
PlainFact is a high-quality human-annotated dataset with fine-grained explanation (i.e., added information) annotations.
1 paper · 0 benchmarks
PreRAID (Prescreening Rheumatoid Arthritis Information Database (PreRAID))
PreRAID is a structured dataset designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in Rheumatoid Arthritis (RA) diagnosis.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.