Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 123 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5857–5904 of 12,172
XLEnt consists of parallel entity mentions in 120 languages aligned with English.
3 papers · 0 benchmarks
We present XHate-999, a multi-domain and multilingual evaluation data set for abusive language detection.
3 papers · 0 benchmarks
YTD-18M is a large-scale corpus of 18M video-based dialogues, constructed from web videos: crucial to the data collection pipeline is a pretrained language model that converts error-prone automatic transcripts to a cleaner dialogue format…
3 papers · 0 benchmarks
Yahoo S5 (Yahoo S5 - A Labeled Anomaly Detection Dataset)
Automatic anomaly detection is critical in today's world where the sheer volume of data makes it impossible to tag outliers manually.
3 papers · 0 benchmarks
Youtbean is a dataset created from closed captions of YouTube product review videos.
3 papers · 0 benchmarks
A new English language dataset structured for task-oriented evaluation on unseen tasks.
3 papers · 0 benchmarks
This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages.
3 papers · 0 benchmarks
Approximately 240,000 documents were collected and aligned using the Hunalign tool.
3 papers · 0 benchmarks
characterRelations dataset contains 2,170 annotations of character relations in 109 literary texts, as documented in characterRelations.pdf.
3 papers · 0 benchmarks
This research aimed at the case of customers default payments in Taiwan and compares the predictive accuracy of probability of default among six data mining methods.
3 papers · 0 benchmarks
The ePiC dataset is a unique and high-quality crowdsourced collection of narratives specifically designed for testing abstract language understanding in the context of proverbs.
3 papers · 0 benchmarks
esXNLI is a bilingual NLI dataset.
3 papers · 0 benchmarks
By releasing this dataset, we aim at providing a new testbed for computer vision techniques using Deep Learning.
3 papers · 0 benchmarks
6981 SAT-level geometry problem with complete natural language description, geometric shapes, formal language annotations, and theorem sequences annotations.
3 papers · 0 benchmarks
iFakeFaceDB is a face image dataset for the study of synthetic face manipulation detection, comprising about 87,000 synthetic face images generated by the Style-GAN model and transformed with the GANprintR approach.
3 papers · 0 benchmarks
iWildCam 2021 is a dataset for counting the number of animals of each species that appear in sequences of images captured with camera traps.
3 papers · 0 benchmarks
A large, crowd-sourced dataset for the Native Language Identification (NLI) task.
3 papers · 1 benchmark
This package provides utilities for generation, filtering, solving, visualizing, and processing of mazes for training ML systems.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
putEMG and putEMG-Force datasets are databases of surface electromyographic activity recorded from forearm.
3 papers · 0 benchmarks
rSoccer is an open-source simulator for the IEEE Very Small Size Soccer and the Small Size League optimized for reinforcement learning experiments.
3 papers · 0 benchmarks
collected by one VLP-16 in a small vehicle (1m x 1m)
3 papers · 1 benchmark
We provide two image stacks where each contains 20 sections from serial section Transmission Electron Microscopy (ssTEM) of the Drosophila melanogaster third instar larva ventral nerve cord.
3 papers · 1 benchmark
Multi-view image dataset of seven objects under indoor lighting, for the purpose of multi-view 3D reconstruction and inverse rendering.
3 papers · 0 benchmarks
xMIND (A Multilingual Dataset for Cross-lingual News Recommendation)
xMIND is an open, large-scale multilingual news dataset for multi- and cross-lingual news recommendation.
3 papers · 0 benchmarks
This 27 Class American Sign Language-based dataset consists of photographs collected from 173 individuals asked to display gestures with their hands.
2 papers · 0 benchmarks
Contains 10⁷ points, sampled from 20 clusters, with incremental concept drift - On each batch (of size 1000) the mean of each of the clusters moves a random (small) length in some random direction, the means move independently of each…
2 papers · 0 benchmarks
The dataset is a .h5 file comprised of entries with keys of the form (n,m), denoting the dimensions of the system matrix on which the simulations have been performed.
2 papers · 0 benchmarks
30MQA (30M Factoid Question-Answer Corpus)
An enormous question answer pair corpus produced by applying a novel neural network architecture on the knowledge base Freebase to transduce facts into natural language questions.
2 papers · 0 benchmarks
360-SOD contains 500 high-resolution equirectangular images.
2 papers · 0 benchmarks
This work was undertaken by members of the Lincoln Centre for Autonomous Systems, University of Lincoln, UK.
2 papers · 0 benchmarks
3D FRONT HUMAN is a dataset that extends the large-scale synthetic scene dataset 3D-FRONT.
2 papers · 0 benchmarks
The platelet-em dataset contains two 3D scanning electron microscope (EM) images of human platelets, as well as instance and semantic segmentations of those two image volumes.
2 papers · 2 benchmarks
Depth vision has been recently used in many locomotion devices with the objective to ease the life of disabled people toward reaching more ecological lifestyle.
2 papers · 0 benchmarks
3DYoga90 (3DYoga90: A Hierarchical Video Dataset for Yoga Pose Understanding)
3DYoga90 is organized within a three-level label hierarchy.
2 papers · 0 benchmarks
A multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from short-video social media platform - Moj.
2 papers · 0 benchmarks
4D Light Field Dataset is a light field benchmark consisting of 24 carefully designed synthetic, densely sampled 4D light fields with highly accurate disparity ground truth.
2 papers · 1 benchmark
This Kaggle repository is still under construction (as of October 2022).
2 papers · 0 benchmarks
This is 4 sketch style (4SKST) dataset, from the research paper "Semi-supervised reference-based sketch extraction using a contrastive learning framework" Dataset consists one of four different styles of sketches paired to color images.
2 papers · 0 benchmarks
Description: 5,011 Images – Human Frontal face Data (Male).
2 papers · 0 benchmarks
We crawled 5000 paper, slide pairs from conference proceeding websites.
2 papers · 0 benchmarks
This is the dataset which contains the ' limitation' text from all papers of ACL 2023
2 papers · 0 benchmarks
ADORE (A benchmark dataset for machine learning in ecotoxicology)
ADORE is a benchmark dataset for machine learning for ecotixicology, covering acute aquatic toxicity in three relevant taxonomic groups (fish, crustaceans, and algae).
2 papers · 1 benchmark
The ADP dataset consists of over 200,000 experimental crystal structures curated from the Cambridge Structural Database (CSD).
2 papers · 0 benchmarks
AG-ReID.v2 (Aerial-Ground Person Re-identification)
Aerial-ground person re-identification (Re-ID) presents unique challenges in computer vision, stemming from the distinct differences in viewpoints, poses, and resolutions between high-altitude aerial and ground-based cameras.
2 papers · 1 benchmark
Antonio Gulli’s corpus of news articles is a collection of more than 1 million news articles.
2 papers · 0 benchmarks
AI-ArtBench: An AI-generated Artistic Dataset AI-ArtBench is a dataset that contains 180,000+ art images.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.