Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 232 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11089–11136 of 12,172
The dataset was curated from the 1% data sample file of the Wikipedia-based Image Text (WIT) Dataset.
1 paper · 0 benchmarks
Purpose: Patients with 1p/19q codeleted low-grade glioma (LGG) have longer overall survival and better treatment response than patients with 1p/19q intact tumors.
1 paper · 0 benchmarks
gtzanmusicspeech is a dataset for music/speech discrimination.
1 paper · 0 benchmarks
hERG is a large-scale biophysics federated molecular dataset related to cardiac toxicity.
1 paper · 0 benchmarks
This dataset of hand drawn images of molecular depictions incorporates 4 sub dataset: dataset to train atom-type object detection model dataset to train bond-type object detection model dataset to train charge-type object detection model…
1 paper · 0 benchmarks
Nutrient loadings and boundary conditions required by the FICOS model.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CoCoPops is a meta-corpus of melodic and harmonic transcriptions of popular music.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The i3-video dataset contains "is-it-instructional" annotations for 6.4k videos from Youtube-8M.
1 paper · 0 benchmarks
A new dataset consisting of 64 people with different expressions and hairstyles.
1 paper · 0 benchmarks
Online web communities often face bans for violating platform policies, encouraging their migration to alternative platforms.
1 paper · 0 benchmarks
Fashion trends are constantly evolving, but a trained eye can estimate with some accuracy the signature elements of a particular designer's style.
1 paper · 2 benchmarks
iFLYTEK and ChangGuang Satellite jointly held the challenge of extracting cultivated land from high-resolution remote sensing images.
1 paper · 0 benchmarks
iLur News Texts is a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society.
1 paper · 0 benchmarks
A dataset that consists of 20 actions of various actors, such as tennis serves, yoga and Tai Chi.
1 paper · 0 benchmarks
The iNaturalist Fine-Grained Geolocation dataset is an extension of the iNaturalist dataset with complementary geolocation information.
1 paper · 0 benchmarks
A dataset consisting of paired ground-level images of species in the inat-2021 dataset and corresponding satellite imagery.
1 paper · 0 benchmarks
iV2V and iV2I+ (AI4Mobile Industrial Wireless Datasets: iV2V and iV2I+)
This dataset provides wireless measurements from two industrial testbeds: iV2V (industrial Vehicle-to-Vehicle) and iV2I+ (industrial Vehicular-to-Infrastructure plus sensor).
1 paper · 0 benchmarks
Easily generate simple continual learning benchmarks.
1 paper · 0 benchmarks
Please see our website and code repository for detailed description.
1 paper · 0 benchmarks
A dataset for Image-Goal Navigation in Habitat based on Gibson scenes.
1 paper · 0 benchmarks
inaGVAD (InaGVAD : a Challenging French TV and Radio Corpus annotated for Voice Activity Detection and Speaker Gender Segmentation)
InaGVAD is a Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) dataset designed for representing the acoustic diversity of French TV and Radio programs.
1 paper · 0 benchmarks
IndirectRequests is an LLM-generated dataset of user utterances in a task-oriented dialogue setting where the user does not directly specify their preferred slot value.
1 paper · 0 benchmarks
"The Chicago Face Database was developed at the University of Chicago by Debbie S.
1 paper · 0 benchmarks
We release the datasets to replicate the results of Coordinated Reply Attacks in Influence Operations: Characterization and Detection'.
1 paper · 0 benchmarks
jaCappella is a corpus of Japanese a cappella vocal ensembles (jaCappella corpus) for vocal ensemble separation and synthesis.
1 paper · 0 benchmarks
The presented dataset contains 10,000 Jupyter notebooks, each of which contains at least one error.
1 paper · 0 benchmarks
It is a competition on kaggle with stroke Prediction, which is heavily imbalanced.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.