Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 113 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5377–5424 of 12,172
HeLa cells stably expressing H2b-GFP Mitocheck Consortium
3 papers · 2 benchmarks
FoodLogoDet-1500 is a new large-scale publicly available food logo dataset, which has 1,500 categories, about 100,000 images and about 150,000 manually annotated food logo objects.
3 papers · 0 benchmarks
Ford Campus Vision and Lidar Data Set is a dataset collected by an autonomous ground vehicle testbed, based upon a modified Ford F-250 pickup truck.
3 papers · 0 benchmarks
The Forms Dataset is a dataset for document structure extraction comprising of 5K forms.
3 papers · 0 benchmarks
Fraxtil is an audio dataset where given a raw audio track, the goal is to produce a choreography step chart, similar to those used in the Dance Dance Revolution video game.
3 papers · 0 benchmarks
FreSaDa is a French satire dataset for cross-domain satire detection, which is composed of 11,570 articles from the news domain.
3 papers · 0 benchmarks
FunQA is a challenging video question answering (QA) dataset specifically designed to evaluate and enhance the depth of video reasoning based on counter-intuitive and fun videos.
3 papers · 0 benchmarks
We present a new large-scale photorealistic panoramic dataset named FutureHouse, which has the following characteristics.
3 papers · 0 benchmarks
GAS (Grasp Area Segmentation)
GAS (Grasp Area Segmentation) dataset consists of 10089 RGB images of cluttered scenes grouped into 1121 grasp-area segmentation tasks.
3 papers · 0 benchmarks
These images were generated using UnityEyes simulator, after including essential eyeball physiology elements and modeling binocular vision dynamics.
3 papers · 0 benchmarks
The GDELT Project is a remarkable initiative that monitors our world by analyzing global news from various sources.
3 papers · 1 benchmark
GENIE (GENeratIve Evaluation)
GENIE, which stands for GENeratIve Evaluation, is a system designed to standardize human evaluations across different text generation tasks.
3 papers · 0 benchmarks
GESTURES (Generalized Epileptic Seizure classification from video-Telemetry Using REcurrent convolutional neural networkS)
This is the dataset to support the paper: Fernando Pérez-García et al., 2021, Transfer Learning of Deep Spatiotemporal Networks to Model Arbitrarily Long Videos of Seizures.
3 papers · 0 benchmarks
Most publications that aim to optimize neural networks for CBIR, train and test their models on domain specific datasets.
3 papers · 0 benchmarks
GarmentCodeData (GarmentCodeData: A Dataset of 3D Made-to-Measure Garments With Sewing Patterns)
GarmentCodeData contains 115,000 data points that cover a variety of designs in many common garment categories: tops, shirts, dresses, jumpsuits, skirts, pants, etc., fitted to a variety of body shapes sampled from a custom statistical…
3 papers · 0 benchmarks
A high-quality dataset for machine translation evaluation that aims at being one of the first non-synthetic gender-balanced test datasets.
3 papers · 0 benchmarks
This dataset refers to the two images acquired by the GeoEye-1 satellite, representing London and Trenton, respectively.
3 papers · 1 benchmark
GeoGLUE (GeoGraphic Language Understanding Evaluation Benchmark)
GeoGLUE is a GeoGraphic Language Understanding Evaluation benchmark, which consists of six geographic text-related tasks, including geographic textual similarity on recall, geotagged geographic elements tagging, geographic composition…
3 papers · 0 benchmarks
Geoclidean-Constraints dataset consists of 20 concepts and 40 tasks, created from permutations of line and circle construction rules with various constraints describing the relationship between objects.
3 papers · 0 benchmarks
Giantsteps is a dataset that includes songs in major and minor scales for all pitch classes, i.e., a 24-way classification task.
3 papers · 0 benchmarks
Are you the kind of person who makes a lot of typos when writing code?
3 papers · 0 benchmarks
The GitTables-SemTab dataset is a subset of the GitTables dataset and was created to be used during the SemTab challenge.
3 papers · 2 benchmarks
The GlassTemp dataset is collected from Polyinfo.
3 papers · 1 benchmark
GlotScript-R is a resource that provides the attested writing systems for more than 7,000 languages.
3 papers · 0 benchmarks
Goldfinch is a dataset for fine-grained recognition challenges.
3 papers · 0 benchmarks
This noisy speech test set is created from the Google Speech Commands v2 [1] and the Musan dataset[2].
3 papers · 1 benchmark
GraSP (Holistic and Multi-Granular Surgical Scene Understanding of Prostatectomies)
Holistic and Multi-Granular Surgical Scene Understanding of Prostatectomies (GraSP) dataset, a curated benchmark that models surgical scene understanding as a hierarchy of complementary tasks with varying levels of granularity.
3 papers · 1 benchmark
GroOT (Grounded Multiple Object Tracking)
One of the recent trends in vision problems is to use natural language captions to describe the objects of interest.
3 papers · 0 benchmarks
HASY is a dataset of single symbols similar to MNIST.
3 papers · 0 benchmarks
The dataset consists of the features associated with 402 5-second sound samples.
3 papers · 0 benchmarks
An autnonomous driving dataset and benchmark for optical flow.
3 papers · 0 benchmarks
HDM05 is a MoCap (motion capture) dataset.
3 papers · 1 benchmark
HELOC (Home Equity Line of Credit)
HELOC The HELOC dataset from FICO.
3 papers · 1 benchmark
HOPE-Image (Household Objects for Pose Estimation)
The NVIDIA HOPE datasets consist of RGBD images and video sequences with labeled 6-DoF poses for 28 toy grocery objects.
3 papers · 0 benchmarks
HOPE-Video (Household Objects for Pose Estimation)
The HOPE-Video dataset contains 10 video sequences (2038 frames) with 5-20 objects on a tabletop scene captured by a robot arm-mounted RealSense D415 RGBD camera.
3 papers · 0 benchmarks
Multilingual text collection extracted from the Internet Archive and Common Crawl archives.
3 papers · 0 benchmarks
HSD (Honda Scenes Dataset)
An annotated dataset is released to enable dynamic scene classification that includes 80 hours of diverse high quality driving video data clips collected in the San Francisco Bay area.
3 papers · 0 benchmarks
HSI-Drive is the hyperspectral image (HSI) dataset created by the Digital Electronics Design Group (GDED) of the University of the Basque Country (UPV/EHU).
3 papers · 1 benchmark
HSPACE (Human-SPACE) is a large-scale photo-realistic dataset of animated humans placed in complex synthetic indoor and outdoor environments.
3 papers · 1 benchmark
The data set contains several speakers.
3 papers · 2 benchmarks
HUME-VB (The Hume Vocal Bursts Dataset)
The Hume Vocal Burst Database (H-VB) includes all train, validation, and test recordings and corresponding emotion ratings for the train and validation recordings.
3 papers · 7 benchmarks
HaVG (Hausa Visual Genome Dataset)
A dataset that contains the description of an image or a section within the image in Hausa and its equivalent in English.
3 papers · 0 benchmarks
This dataset is built from Twitter and contains 1290 hate tweet and counterspeech reply pairs.
3 papers · 0 benchmarks
HateScore (HateScore : Human-in-the-Loop and Neutral Korean Multi-label Online Hate Speech Dataset)
2.2K neutral sentences from Wikipedia 1.7K additionally labeled sentences generated by the Human-in-the-Loop procedure (based on Korean Unsmile Dataset Base Model) 7.1K rule-generated neutral sentences
3 papers · 0 benchmarks
HatemojiCheck is a test suite for detecting emoji-based hate of 3,930 test cases covering seven functionalities of emoji-based hate and six identities.
3 papers · 0 benchmarks
Hazards&Robots (Hazards&Robots: A Dataset for Visual Anomaly Detection in Robotics)
We consider the problem of detecting, in the visual sensing data stream of an autonomous mobile robot, semantic patterns that are unusual (i.e., anomalous) with respect to the robot’s previous experience in similar environments.
3 papers · 0 benchmarks
HeadlineCause is a dataset for detecting implicit causal relations between pairs of news headlines.
3 papers · 0 benchmarks
Healthline is a nutrition related dataset for multi-document summarization, using scientific studies.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.