Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 37 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 1729–1776 of 12,172
iNat2021 is a large-scale image dataset collected and annotated by community scientists that contains over 2.7M images from 10k different species.
27 papers · 0 benchmarks
We have created three new Reading Comprehension datasets constructed using an adversarial model-in-the-loop.
26 papers · 2 benchmarks
AlignBench is a comprehensive benchmark designed specifically for evaluating the alignment performance of large Chinese language models (LLMs).
26 papers · 0 benchmarks
Animal Kingdom is a large and diverse dataset that provides multiple annotated tasks to enable a more thorough understanding of natural animal behaviors.
26 papers · 2 benchmarks
CRVD (Captured Raw Video Denoising)
The CRVD dataset consists of 55 groups of noisy-clean videos with ISO values ranging from 1600 to 25600.
26 papers · 1 benchmark
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
Chart-to-text is a large-scale benchmark with two datasets and a total of 44,096 charts covering a wide range of topics and chart types.
26 papers · 0 benchmarks
Cityscapes-VPS is a video extension of the Cityscapes validation split.
26 papers · 1 benchmark
ContactPose is a dataset of hand-object contact paired with hand pose, object pose, and RGB-D images.
26 papers · 1 benchmark
DALES (DALES: A Large-scale Aerial LiDAR Data Set for Semantic Segmentation)
We present the Dayton Annotated LiDAR Earth Scan (DALES) data set, a new large-scale aerial LiDAR data set with over a half-billion hand-labeled points spanning 10 square kilometers of area and eight object categories.
26 papers · 2 benchmarks
Using the validation set (100 images) from the widely used DIV2K dataset, we blurred and subsampled each image with a different, randomly generated kernel.
26 papers · 2 benchmarks
DUDE (Document UnderstanDing of Everything)
DUDE is formulated as an instance of Document Question Answering (DocQA) to evaluate how well current solutions deal with multi-page documents, if they can navigate and reason over the layout, and if they can generalize these skills to…
26 papers · 0 benchmarks
The Drive&Act dataset is a state of the art multi modal benchmark for driver behavior recognition.
26 papers · 1 benchmark
FairytaleQA is a dataset focusing on narrative comprehension of kindergarten to eighth-grade students.
26 papers · 2 benchmarks
Contains 1024 pairs of high-quality images and covers diverse scenarios.
26 papers · 0 benchmarks
Object detection benchmark for logo detection.
26 papers · 3 benchmarks
The GenericsKB contains 3.4M+ generic sentences about the world, i.e., sentences expressing general truths such as "Dogs bark," and "Trees remove carbon dioxide from the atmosphere." Generics are potentially useful as a knowledge source…
26 papers · 0 benchmarks
Head and Neck Tumor Segmentation
26 papers · 0 benchmarks
Harm-C is a dataset for detecting harmful memes related to Covid-19.
26 papers · 0 benchmarks
HomebrewedDB is a dataset for 6D pose estimation mainly targeting training from 3D models (both textured and textureless), scalability, occlusions, and changes in light conditions and object appearance.
26 papers · 0 benchmarks
HumanEvalPack is an extension of OpenAI's HumanEval to cover 6 total languages across 3 tasks.
26 papers · 1 benchmark
IIRC (Incomplete Information Reading Comprehension)
Contains more than 13K questions over paragraphs from English Wikipedia that provide only partial information to answer them, with the missing information occurring in one or more linked documents.
26 papers · 0 benchmarks
ImageNet VID is a large-scale public dataset for video object detection and contains more than 1M frames for training and more than 100k frames for validation.
26 papers · 1 benchmark
A benchmark dataset for out-of-distribution detection.
26 papers · 1 benchmark
M2DGR (a Multi-modal and Multi-scenario SLAM Dataset for Ground Robots)
We collected long-term challenging sequences for ground robots both indoors and outdoors with a complete sensor suite, which includes six surround-view fish-eye cameras, a sky-pointing fish-eye camera, a perspective color camera, an event…
26 papers · 0 benchmarks
MNIST8M is derived from the MNIST dataset by applying random deformations and translations to the dataset.
26 papers · 0 benchmarks
Our dataset was made of videos from MSU Video Upscalers Benchmark Dataset, MSU Video Super-Resolution Benchmark Dataset and MSU Super-Resolution for Video Compression Benchmark Dataset.
26 papers · 1 benchmark
MiniHack is a sandbox framework for easily designing rich and diverse environments for Reinforcement Learning (RL).
26 papers · 0 benchmarks
Dataset is constructed from single intent dataset SNIPS.
26 papers · 2 benchmarks
MultiDoc2Dial (MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents)
MultiDoc2Dial is a new task and dataset on modeling goal-oriented dialogues grounded in multiple documents.
26 papers · 0 benchmarks
MultiviewX is a synthetic Multiview pedestrian detection dataset.
26 papers · 2 benchmarks
PeMSD7 is traffic data in District 7 of California consisting of the traffic speed of 228 sensors while the period is from May to June in 2012 (only weekdays) with a time interval of 5 minutes.
26 papers · 2 benchmarks
PixelHelp includes 187 multi-step instructions of 4 task categories deined in https://support.google.com/pixelphone and annotated by human.
26 papers · 0 benchmarks
The Q-Bench includes three realms for low-level vision: perception (A1), description (A2), and assessment (A3).
26 papers · 0 benchmarks
The exact pre-processing steps used to construct the MNIST dataset have long been lost.
26 papers · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
26 papers · 1 benchmark
The SUN09 dataset consists of 12,000 annotated images with more than 200 object categories.
26 papers · 0 benchmarks
Screen2Words is a large-scale screen summarization dataset annotated by human workers.
26 papers · 0 benchmarks
The PASCAL-Scribble Dataset is an extension of the PASCAL dataset with scribble annotations for semantic segmentation.
26 papers · 0 benchmarks
🤖 Robo3D - The SemanticKITTI-C Benchmark SemanticKITTI-C is an evaluation benchmark heading toward robust and reliable 3D semantic segmentation in autonomous driving.
26 papers · 1 benchmark
The SentiCap dataset contains several thousand images with captions with positive and negative sentiments.
26 papers · 0 benchmarks
The StarCraft II Learning Environment (S2LE) is a reinforcement learning environment based on the game StarCraft II.
26 papers · 0 benchmarks
Street Scene is a dataset for video anomaly detection.
26 papers · 3 benchmarks
UrbanCars facilitates multi-shortcut learning under the controlled setting with two shortcuts—background and co-occurring object.
26 papers · 0 benchmarks
UrbanScene3D is a large scale urban scene dataset associated with a handy simulator based on Unreal Engine 4 and AirSim, which consists of both man-made and real-world reconstruction scenes in different scales, referred to as UrbanScene3D.
26 papers · 0 benchmarks
VALSE (VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena)
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic…
26 papers · 12 benchmarks
VAW (Visual Attributes in the Wild)
VAW is a large scale visual attributes dataset with explicitly labelled positive and negative attributes.
26 papers · 0 benchmarks
VOID (Visual Odometry with Inertial and Depth)
The dataset was collected using the Intel RealSense D435i camera, which was configured to produce synchronized accelerometer and gyroscope measurements at 400 Hz, along with synchronized VGA-size (640 x 480) RGB and depth streams at 30 Hz.
26 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.