Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 75 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3553–3600 of 12,172
We manually performed the task of Open Information Extraction on 5 short documents, elaborating tentative guidelines for the task, and resulting in a ground truth reference of 347 tuples.
8 papers · 1 benchmark
Wiki (Web Traffic Time Series Forecasting)
Context There's a story behind every dataset and here's your opportunity to share yours.
8 papers · 3 benchmarks
WikiTableT contains Wikipedia article sections and their corresponding tabular data and various metadata.
8 papers · 0 benchmarks
WildReceipt is a collection of receipts.
8 papers · 0 benchmarks
Winogender Schemas is a novel, Winograd schema-style set of minimal pair sentences that differ only by pronoun gender.
8 papers · 0 benchmarks
WorldStrat (The WorldStrat Dataset: Open High-Resolution Satellite Imagery With Paired Multi-Temporal Low-Resolution)
Nearly 10,000 km² of free high-resolution and paired multi-temporal low-resolution satellite imagery of unique locations which ensure stratified representation of all types of land-use across the world: from agriculture to ice caps, from…
8 papers · 0 benchmarks
X-Humans consists of 20 subjects (11 males, 9 females) with various clothing types and hair style.
8 papers · 0 benchmarks
X3D is a dataset containing 15 scenes and covering 4 applications for X-ray 3D reconstruction.
8 papers · 2 benchmarks
XAI-Bench is a suite of synthetic datasets along with a library for benchmarking feature attribution algorithms.
8 papers · 0 benchmarks
The YouTube-100M data set consists of 100 million YouTube videos: 70M training videos, 10M evaluation videos, and 20M validation videos.
8 papers · 0 benchmarks
YouTube-ASL is a large-scale, open-domain corpus of American Sign Language (ASL) videos and accompanying English captions drawn from YouTube.
8 papers · 0 benchmarks
We address the problem of automatically learning the main steps to complete a certain task, such as changing a car tire, from a set of narrated instruction videos.
8 papers · 2 benchmarks
ZEB (Zero-shot Evaluation Benchmark)
A evaluation benchmark ZEB for image matching by merging 8 real-world datasets and 4 simulated datasets with diverse image resolutions, scene conditions and view points.
8 papers · 1 benchmark
LightStage is a multi-view dataset, which is proposed in NeuralBody.
8 papers · 0 benchmarks
Research on semantic segmentation of traffic scenes using color and polarization information (including training and testing sets).
8 papers · 1 benchmark
https://github.com/YatingMusic/compound-word-transformer
8 papers · 0 benchmarks
Dataset of high-pT jets from simulations of LHC proton-proton collisions Prepared for FastML/HLS4ML studies: https://fastmachinelearning.org Includes: High level features (see https://arxiv.org/abs/1804.06913) Images: jet images with up to…
8 papers · 0 benchmarks
The data is provided by 10x Genomics under "Single Cell 3' Paper: Zheng et al.
7 papers · 0 benchmarks
3D-FRONT (3D Furnished Rooms with layOuts and semaNTics) is large-scale, and comprehensive repository of synthetic indoor scenes highlighted by professionally designed layouts and a large number of rooms populated by high-quality textured…
7 papers · 0 benchmarks
3DOH50K is the first real 3D human dataset for the problem of human reconstruction and pose estimation in occlusion scenarios.
7 papers · 1 benchmark
The 7-Scenes dataset is a collection of tracked RGB-D camera frames.
7 papers · 0 benchmarks
ACAV100M (Automatically Curated Audio-Visual)
ACAV100M processes 140 million full-length videos (total duration 1,030 years) which are used to produce a dataset of 100 million 10-second clips (31 years) with high audio-visual correspondence.
7 papers · 0 benchmarks
ACES (A Translation Accuracy Challenge Set)
ACES a dataset consisting of 68 phenomena ranging from simple perturbations at the word/character level to more complex errors based on discourse and real-world knowledge.
7 papers · 1 benchmark
AGAR (Annotated Germs for Automated Recognition)
The Annotated Germs for Automated Recognition (AGAR) dataset is an image database of microbial colonies cultured on an agar plate.
7 papers · 0 benchmarks
AIOZ-GDANCE comprises 16.7 hours of whole-body motion and music audio of group dancing.
7 papers · 1 benchmark
AVisT (A Benchmark for Visual Object Tracking in Adverse Visibility)
One of the key factors behind the recent success in visual tracking is the availability of dedicated benchmarks.
7 papers · 1 benchmark
Aachen Day-Night v1.1 dataset is an extended version of the original Aachen Day-Night dataset.
7 papers · 1 benchmark
We introduce ArtBench-10, the first class-balanced, high-quality, cleanly annotated, and standardized dataset for benchmarking artwork generation.
7 papers · 1 benchmark
ArtiFact (Artificial and Factual Image Dataset for Synthetic Image Detection)
The ArtiFact dataset is a large-scale image dataset that aims to include a diverse collection of real and synthetic images from multiple categories, including Human/Human Faces, Animal/Animal Faces, Places, Vehicles, Art, and many other…
7 papers · 0 benchmarks
The Atari Grand Challenge dataset is a large dataset of human Atari 2600 replays.
7 papers · 0 benchmarks
BANDON is a dataset for building change detection with off-nadir aerial images dataset, which is composed of off-Nadir image pairs of urban and rural areas.
7 papers · 0 benchmarks
Prediction of Finger Flexion IV Brain-Computer Interface Data Competition The goal of this dataset is to predict the flexion of individual fingers from signals recorded from the surface of the brain (electrocorticography (ECoG)).
7 papers · 1 benchmark
BOVText is a new large-scale benchmark dataset named Bilingual, Open World Video Text(BOVText), the first large-scale and multilingual benchmark for video text spotting in a variety of scenarios.
7 papers · 0 benchmarks
BS-RSC is a real-world rolling shutter (RS) correction dataset and a corresponding model to correct the RS frames in a distorted video.
7 papers · 1 benchmark
Backstabber’s Knife Collection is a dataset of 174 malicious software packages that were used in real-world attacks on open source software supply chains, and which were distributed via the popular package repositories npm, PyPI, and…
7 papers · 0 benchmarks
Bamboo Dataset is a mega-scale and information-dense dataset for both classification and detection pre-training.
7 papers · 0 benchmarks
An annotated dataset of ~50K news that can be used for building automated fake news detection systems for a low resource language like Bangla.
7 papers · 0 benchmarks
BiSECT is a dataset for sentence simplification, which is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary.
7 papers · 0 benchmarks
BigDetection is a new large-scale benchmark to build more general and powerful object detection systems.
7 papers · 1 benchmark
BinKit is a binary code similarity analysis (BCSA) benchmark.
7 papers · 0 benchmarks
Biwi Kinect Head Pose is a challenging dataset mainly inspired by the automotive setup.
7 papers · 0 benchmarks
A new challenging dataset that can be used for many pattern recognition tasks.
7 papers · 1 benchmark
BraTS 2020 (RSNA-ASNR-MICCAI Brain Tumor Segmentation BraTS Challenge)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
7 papers · 0 benchmarks
This dataset is a combination of the following three datasets : figshare, SARTAJ dataset and Br35H This dataset contains 7022 images of human brain MRI images which are classified into 4 classes: glioma - meningioma - no tumor and…
7 papers · 3 benchmarks
CARL (Context Adaptive RL)
CARL (context adaptive RL) provides highly configurable contextual extensions to several well-known RL environments.
7 papers · 1 benchmark
CASME II (Chinese Academy of Sciences Micro-Expression II)
The Chinese Academy of Sciences Micro-Expression dataset (CASME II) consists of 255 videos, elicited from 26 participants.
7 papers · 1 benchmark
CEFR-SP contains 17k English sentences annotated with the levels based on the Common European Framework of Reference for Languages assigned by English-education professionals.
7 papers · 0 benchmarks
CHAD (Charlotte Anomaly Dataset)
CHAD: Charlotte Anomaly Dataset CHAD is high-resolution, multi-camera dataset for surveillance video anomaly detection.
7 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.