Home › Datasets › modality › Biology

Biology datasets

archive 2025-07-28

70 datasets carry the modality tag "Biology", ordered by the archive's paper count. Page 1 of 2: 48 shown of 70. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Biology datasets 1–48 of 70

PROTEINS is a dataset of proteins that are classified as enzymes or non-enzymes.
371 papers · 1 benchmark
Contains hundreds of frontal view X-rays and is the largest public resource for COVID-19 image and prognostic data, making it a necessary resource to develop and evaluate tools to aid in the treatment of COVID-19.
35 papers · 1 benchmark
MHIST (Minimalist Histopathology image analysis dataset)
The minimalist histopathology image analysis dataset (MHIST) is a binary classification dataset of 3,152 fixed-size images of colorectal polyps, each with a gold-standard label determined by the majority vote of seven board-certified…
28 papers · 1 benchmark
LIVECell (Label-free In Vitro image Examples of Cells)
The LIVECell (Label-free In Vitro image Examples of Cells) dataset is a large-scale microscopic image dataset for instance-segmentation of individual cells in 2D cell cultures.
18 papers · 1 benchmark
Yeast dataset consists of a protein-protein interaction network.
18 papers · 0 benchmarks
FLIP (Fitness Landscape Inference for Proteins)
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function).
11 papers · 0 benchmarks
MIPE (Improving Paratope and Epitope Prediction by Multi-Modal Contrastive Learning and Interaction Informativeness Estimation)
Datasets.
7 papers · 1 benchmark
2D HeLa is a dataset of fluorescence microscopy images of HeLa cells stained with various organelle-specific fluorescent dyes.
6 papers · 0 benchmarks
RxRx1 is a biological dataset designed specifically for the systematic study of batch effect correction methods.
6 papers · 1 benchmark
BB-norm-habitat (Bacteria Biotope - entity normalization - bacterial habitat)
In the BB-norm modality of this task, participant systems had to normalize textual entity mentions according to the OntoBiotope ontology for habitats.
5 papers · 0 benchmarks
BB-norm-phenotype (Bacteria Biotope - entity normalization - phenotype)
In the BB-norm modality of this task, participant systems had to normalize textual entity mentions according to the OntoBiotope ontology for phenotypes.
5 papers · 0 benchmarks
CBC (Complete Blood Count)
The complete blood count (CBC) dataset contains 360 blood smear images along with their annotation files splitting into Training, Testing, and Validation sets.
5 papers · 0 benchmarks
NucMM is a dataset for segmenting 3D cell nuclei from microscopy image volumes that pushes the task forward to the sub-cubic millimeter scale.
5 papers · 0 benchmarks
PECAN (Paratope-Epitope Complexes for Antibody Networks (PECAN))
The PECAN dataset provides structural data for antibody-antigen interactions, specifically curated for paratope and epitope binding site prediction.
5 papers · 1 benchmark
As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, we present the BIOSCAN-5M Insect dataset to the machine learning community.
4 papers · 0 benchmarks
CausalBench is a comprehensive benchmark suite for evaluating network inference methods on large-scale perturbational single-cell gene expression data.
4 papers · 0 benchmarks
MassSpecGym (MassSpecGym: A benchmark for the discovery and identification of molecules)
MassSpecGym provides three challenges for benchmarking the discovery and identification of new molecules from MS/MS spectra: - 💥 De novo molecule generation (MS/MS spectrum → molecular structure) - ✨ Bonus chemical formulae challenge…
4 papers · 6 benchmarks
For each dataset we provide a short description as well as some characterization metrics.
4 papers · 0 benchmarks
The dataset is designed specifically to solve a range of computer vision problems (2D-3D tracking, posture) faced by biologists while designing behavior studies with animals.
3 papers · 0 benchmarks
CREMP is a resource generated for the rapid development and evaluation of machine learning models for macrocyclic peptides.
3 papers · 0 benchmarks
FOBIE (Focused Open Biological Information Extraction)
The Focused Open Biology Information Extraction (FOBIE) dataset aims to support IE from Computer-Aided Biomimetics.
3 papers · 0 benchmarks
PWDB (Pulse Wave Database)
Overview This database of simulated arterial pulse waves is designed to be representative of a sample of pulse waves measured from healthy adults.
3 papers · 0 benchmarks
fluocells (Fluorescent Neuronal Cells)
By releasing this dataset, we aim at providing a new testbed for computer vision techniques using Deep Learning.
3 papers · 0 benchmarks
This work was undertaken by members of the Lincoln Centre for Autonomous Systems, University of Lincoln, UK.
2 papers · 0 benchmarks
3D Platelet EM (Platelet Electron Microscopy)
The platelet-em dataset contains two 3D scanning electron microscope (EM) images of human platelets, as well as instance and semantic segmentations of those two image volumes.
2 papers · 2 benchmarks
ADORE (A benchmark dataset for machine learning in ecotoxicology)
ADORE is a benchmark dataset for machine learning for ecotixicology, covering acute aquatic toxicity in three relevant taxonomic groups (fish, crustaceans, and algae).
2 papers · 1 benchmark
ATUE is an antibody study benchmark with four real-world supervised tasks covering therapeutic antibody engineering, B cell analysis, and antibody discovery.
2 papers · 0 benchmarks
The H01 dataset is a 1.4 petabyte rendering of a small sample of human brain tissue, released by a collaboration between the Lichtman Laboratory at Harvard University and Google.
2 papers · 0 benchmarks
The LeukemiaAttri dataset is a large-scale, multi-domain collection of microscopy images derived from leukemia patient samples, enriched with detailed morphological information.
2 papers · 2 benchmarks
Results of a high-throughput biological assay measuring the stability of proteins https://github.com/Rocklin-Lab/cdna-display-proteolysis-pipeline From the paper "Here we present cDNA display proteolysis, a method for measuring…
2 papers · 0 benchmarks
To take advantage of the ever-increasing amount of structural data now available, we also trained Paragraph on a larger dataset.
2 papers · 1 benchmark
ProteinGym is a collection of benchmarks aiming at comparing the ability of models to predict the effects of protein mutations.
2 papers · 0 benchmarks
The dataset represents data generated from a commonly used model in population genetics.
2 papers · 0 benchmarks
neuronIO (Single cortical neuron (L5PC) input output simulation at 1ms temporal resolution)
Single cortical neurons as deep artificial neural networks This dataset contains training and testing subsets of the input/output relationship of a single cortical layer 5 pyramidal cell (L5PC) neuron at 1ms single spike temporal…
2 papers · 0 benchmarks
In one round of sequencing, 5 fecal pellets from 2 pro-inflammatory environments (Harvard BRI/Johns Hopkins) and 2 pro-survival environments (Broad Institute/Jackson Labs) were sequenced at the 16s rDNA locus.
1 paper · 0 benchmarks
ACCT Data Repository (ACCT is a fast and accessible automatic cell counting tool using machine learning for 2D image segmentation)
This dataset is a collection of fluorescent images from mice in order to test an automatic cell counting tool that we developed.
1 paper · 0 benchmarks
AI-ready multiplex IHC-IF dataset (AI-ready restained and co-registered multiplex dataset for head-and-neck squamous cell carcinoma)
We introduce a new AI-ready computational pathology dataset containing restained and co-registered digitized images from eight head-and-neck squamous cell carcinoma patients.
1 paper · 0 benchmarks
AUR & UMB dataset (Anticancer Efficacy of Auraptene & Umbelliprenin: In Vitro Viability Dataset)
This dataset contains quantitative data on the anticancer effects of the natural coumarins Auraptene (AUR) and Umbelliprenin (UMB) across 27 studies.
1 paper · 1 benchmark
BIRDeep (BIRDeep_AudioAnnotations)
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain.
1 paper · 0 benchmarks
The CATH (Class, Architecture, Topology, Homology) [65] database is a comprehensive resource for protein structure classification that hierarchical group proteins based on their structural features.
1 paper · 1 benchmark
This dataset contains recordings of 32 sound producing insect species with a total 335 files and a length of 57 minutes.
1 paper · 0 benchmarks
The data used for all results in this paper can be found here.
1 paper · 0 benchmarks
ENSeg Dataset Overview This dataset represents an enhanced subset of the ENS dataset.
1 paper · 1 benchmark
The dataset X of this work is an extension of the heartSeg dataset.
1 paper · 1 benchmark
FLIP -- AAV, Designed vs mutant (adeno-associated virus)
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function).
1 paper · 0 benchmarks
Facial Skeletal angles (Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis))
Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis)
1 paper · 0 benchmarks
GO21 is a biomedical knowledge graph that models genes, proteins, drugs, and the hierarchy of the biological processes they participate in.
1 paper · 1 benchmark
This publicly available dataset contains 1613 RGB-D images of field-grown broccoli plants.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.