Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 90 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4273–4320 of 12,172
GOO (Gaze-on-Objects) is a dataset for gaze object prediction, where the goal is to predict a bounding box for a person's gazed-at object.
5 papers · 0 benchmarks
The first AI-generated video detection datasets.
5 papers · 0 benchmarks
The GeoLifeCLEF 2020 dataset is a large-scale remote sensing dataset.
5 papers · 0 benchmarks
Geolife (Geolife GPS Trajectory Dataset)
This is a GPS trajectory dataset collected in (Microsoft Research Asia) GeoLife project by 182 users in a period of over three years (from April 2007 to August 2012).
5 papers · 0 benchmarks
Repair AST parse (syntax) errors in Python code
5 papers · 1 benchmark
GolfDB is a high-quality video dataset created for general recognition applications in the sport of golf, and specifically for the task of golf swing sequencing.
5 papers · 0 benchmarks
The GoogleEarth dataset is collected from Google Earth Studio, including 400 orbit trajectories in Manhattan and Brooklyn.
5 papers · 1 benchmark
We leverage knowledge from foundation models to introduce Grasp-Anything, a new large-scale dataset with 1M (one million) samples and 3M objects, substantially surpassing prior datasets in diversity and magnitude.
5 papers · 0 benchmarks
Grep-BiasIR (Gender Representation-Bias for Information Retrieval)
Grep-BiasIR is a novel thoroughly-audited dataset which aim to facilitate the studies of gender bias in the retrieved results of IR systems.
5 papers · 0 benchmarks
Benchmark for de novo molecular design
5 papers · 0 benchmarks
H-DIBCO 2012 is the International Document Image Binarization Competition which is dedicated to handwritten document images organized in conjunction with ICFHR 2012 conference.
5 papers · 0 benchmarks
HARPER (Exploring 3D Human Pose Estimation and Forecasting from the Robot’s Perspective: The HARPER Dataset)
We introduce HARPER, a novel dataset for 3D body pose estimation and forecast in dyadic interactions between users and \spot, the quadruped robot manufactured by Boston Dynamics.
5 papers · 3 benchmarks
We have added data and cleaned the labels in HC-STVG to build the HC-STVG2.0.
5 papers · 1 benchmark
The dataset consists of 3640 bursts (made up of 28461 images in total), organized into subfolders, plus the results of an image processing pipeline.
5 papers · 0 benchmarks
HDR-GS (HDR-GS: Efficient High Dynamic Range Novel View Synthesis at 1000x Speed via Gaussian Splatting)
This is dataset for high dynamic range novel view synthesis.
5 papers · 1 benchmark
HEV-I (Honda Egocentric View-Intersection Dataset)
Honda Egocentric View-Intersection Dataset (HEV-I) is introduced to enable research on traffic participants interaction modelling, future object localization, as well as learning driver action in challenging driving scenarios.
5 papers · 1 benchmark
HErlev (HErlev Pap Smear Dataset)
5 papers · 1 benchmark
A new RGB-D video dataset, i.e., UCLA Human-Human-Object Interaction (HHOI) dataset, which includes 3 types of human-human interactions, i.e., shake hands, high-five, pull up, and 2 types of human-object-human interactions, i.e., throw and…
5 papers · 0 benchmarks
HJDataset is a large dataset of Historical Japanese Documents with Complex Layouts.
5 papers · 0 benchmarks
HOD (Hand-held Object Dataset)
HOD is a dataset for 3D object reconstruction which contains 35 objects, divided into two subsets named Sculptures and Daily Objects.
5 papers · 1 benchmark
While Video Instance Segmentation (VIS) has seen rapid progress, current approaches struggle to predict high-quality masks with accurate boundary details.
5 papers · 1 benchmark
HaDes is a token-level, reference-free hallucination detection dataset named HAllucination DEtection dataSet.
5 papers · 0 benchmarks
HaN-Seg (The Head and Neck Organ-at-Risk CT & MR Segmentation Challenge)
Cancer in the region of the head and neck (HaN) is one of the most prominent cancers, for which radiotherapy represents an important treatment modality that aims to deliver a high radiation dose to the targeted cancerous cells while…
5 papers · 0 benchmarks
fine-grained location names extraction from disaster-related tweets
5 papers · 1 benchmark
The HateBR dataset is a significant resource for studying offensive language and hate speech detection in Brazilian Portuguese.
5 papers · 0 benchmarks
The Headlines dataset for sarcasm detection is collected from two news website.
5 papers · 0 benchmarks
The Helvipad dataset is a real-world stereo dataset designed for omnidirectional depth estimation.
5 papers · 1 benchmark
A parallel corpus of Hindi and English, and HindMonoCorp, a monolingual corpus of Hindi in their release version 0.5.
5 papers · 0 benchmarks
We introduce HourVideo, a benchmark dataset for hour-long video-language understanding.
5 papers · 0 benchmarks
HowMany-Qa is a object counting dataset.
5 papers · 1 benchmark
A dataset of 69,270,581 video clip, question and answer triplets (v, q, a).
5 papers · 0 benchmarks
Hummingbird is a dataset to examine stylistic lexical cues from human perception and BERT used to characterize their discrepancy.
5 papers · 0 benchmarks
The IAM database contains 13,353 images of handwritten lines of text created by 657 writers.
5 papers · 1 benchmark
IDDA is a large scale, synthetic dataset for semantic segmentation with more than 100 different source visual domains.
5 papers · 0 benchmarks
IG-1B-Targeted is an internal Facebook AI Research dataset that consists of 940 million public images with 1.5K hashtags matching with 1000 ImageNet1K synsets.
5 papers · 0 benchmarks
IGLU is a dataset designed for interactive grounded language understanding.
5 papers · 0 benchmarks
An abnormal activity data-set for research use that contains 4,83,566 annotated frames.
5 papers · 2 benchmarks
IIW (Intrinsic Images in the Wild)
Intrinsic Images in the Wild is a large scale, public dataset for intrinsic image decompositions of real-world scenes selected from the OpenSurfaces dataset.
5 papers · 0 benchmarks
The IPN Hand dataset is a benchmark video dataset with sufficient size, variation, and real-world elements able to train and evaluate deep neural networks for continuous Hand Gesture Recognition (HGR).
5 papers · 0 benchmarks
Consists of user-generated aerial videos from social media with annotations of instance-level building damage masks.
5 papers · 0 benchmarks
The ISIC 2017 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
5 papers · 0 benchmarks
A medical image segmentation challenge at the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2017.
5 papers · 0 benchmarks
A dataset for evaluating text classification, domain adaptation, and active learning models.
5 papers · 0 benchmarks
The image aesthetic benchmark [18] consists of 10800 Flickr photos of four categories, i.e., “animals”, “urban”, “people” and “nature”, and is constructed originally to retrieve beautiful yet unpopular images in social networks.
5 papers · 1 benchmark
ImageNet-Hard is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet).
5 papers · 1 benchmark
ImageNet-Patch: A Dataset for Benchmarking Machine Learning Robustness against Adversarial Patches Adversarial patches are optimized contiguous pixel blocks in an input image that cause a machine-learning model to misclassify it.
5 papers · 0 benchmarks
The ImplicitQA dataset was introduced in the paper ImplicitQA: Going beyond frames towards Implicit Video Reasoning.
5 papers · 1 benchmark
InfiniBench (InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding)
We introduce InfiniBench a comprehensive benchmark for very long video understanding, which presents 1) The longest video duration, averaging 76.34 minutes; 2) The largest number of question-answer pairs, 108.2K; 3) Diversity in questions…
5 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.