Home › Datasets › modality › Biology

Biology datasets

archive 2025-07-28

70 datasets carry the modality tag "Biology", ordered by the archive's paper count. Page 2 of 2: 22 shown of 70. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Biology datasets 49–70 of 70

This data set contains weekly scans of cauliflower and broccoli covering a ten week growth cycle from transplant to harvest.
1 paper · 0 benchmarks
This dataset contains pre-processed versions of datasets introduced in prior works.
1 paper · 0 benchmarks
Marine Microalgae Detection in Microscopy Images dataset contains a total number of images in the dataset is 937 and all the objects in these images were annotated.
1 paper · 0 benchmarks
A dataset of 18,731 proteins with their PDB code, index of the first residue in their respective DSSP file, their residue sequence and 9-category secondary structure sequence (including polyproline helices).
1 paper · 1 benchmark
TPM values together with cell type annotations that were obtained from Alex Pollen on 15/10/15 Source: Low-coverage single-cell mRNA sequencing reveals cellular heterogeneity and activated signaling pathways in developing cerebral cortex
1 paper · 1 benchmark
PubChem18 (PubChem 2018)
A.2.1 AN OPEN, LARGE-SCALE DATASET FOR ZERO-SHOT DRUG DISCOVERY DERIVED FROM PUBCHEM We constructed a large public dataset extracted from PubChem (Kim et al., 2019; Preuer et al., 2018), an open chemistry database, and the largest…
1 paper · 0 benchmarks
3D confocal stacks with corresponding 2D Light-field microscope images Confocal: -Single volume dimension: 1287x1287x64.
1 paper · 0 benchmarks
TCB-DS (Toxigenic Cyanobacteria Dataset)
The TCB-DS dataset is a specialized collection of microscopic images focusing on the automatic recognition of cyanobacteria genera.
1 paper · 0 benchmarks
TERRA-REF (TERRA-REF, An open reference data set from high resolution genomics, phenomics, and imaging sensors)
The ARPA-E funded TERRA-REF project is generating open-access reference datasets for the study of plant sensing, genomics, and phenomics.
1 paper · 0 benchmarks
The EMBO SourceData-NLP dataset (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process.
1 paper · 1 benchmark
VISEM-Tracking is a dataset consisting of 20 video recordings of 30s of spermatozoa with manually annotated bounding-box coordinates and a set of sperm characteristics analyzed by experts in the domain.
1 paper · 0 benchmarks
VesselGraph is a dataset of whole-brain vessel graphs based on specific imaging protocols.
1 paper · 0 benchmarks
YIM Dataset (Yeast Cells in Microstructures Dataset)
An instance segmentation dataset of yeast cells in microstructures.
1 paper · 0 benchmarks
uBench (MicroBench)
Microscopy is a cornerstone of biomedical research, enabling detailed study of biological structures at multiple scales.
1 paper · 0 benchmarks
ALFI (Annotations for Label-Free Images)
ALFI (Annotations for Label-Free Images) is a dataset of images and annotations for label-free microscopy imaging.
0 papers · 0 benchmarks
CAMEO (Continuous automated model evaluation)
Xavier Robin, Juergen Haas, Rafal Gumienny, Anna Smolinski, Gerardo Tauriello, and Torsten Schwede.Continuous automated model evaluation (cameo)—perspectives on the future of fully automated evaluation of structure prediction…
0 papers · 0 benchmarks
This record contains the saddle search output logs for Sella and EON (dimer, with and without GPR acceleration).
0 papers · 0 benchmarks
Genome-wide miRNA detection (Genome-wide hairpins datasets of animals and plants for novel miRNA prediction)
We've made available several genome-wide datasets, which can be used for training microRNA (miRNA) classifiers.
0 papers · 0 benchmarks
The medaka (Oryzias latipes) and the zebrafish (Danio rerio) are used as a model organism for a variety of subjects in biomedical research.
0 papers · 0 benchmarks
SourceData-NLP (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
Introduction: The scientific publishing landscape is expanding rapidly, creating challenges for researchers to stay up-to-date with the evolution of the literature.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
ZooScanNet (ZooScanNet: plankton images captured with the ZooScan)
Plankton was sampled with various nets, from bottom or 500m depth to the surface, in many oceans of the world.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.