Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 59 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2785–2832 of 12,172
MineRL BASALT is an RL competition on solving human-judged tasks.
12 papers · 0 benchmarks
MoVi (Large Multipurpose Motion and Video Dataset)
Contains 60 female and 30 male actors performing a collection of 20 predefined everyday actions and sports movements, and one self-chosen movement.
12 papers · 1 benchmark
MuST-Cinema is a Multilingual Speech-to-Subtitles corpus ideal for building subtitle-oriented machine and speech translation systems.
12 papers · 0 benchmarks
Taxi flow data of New York City with grid 20x10.
12 papers · 1 benchmark
Neptune (Neptune Long Video Understanding Benchmark)
Neptune is a dataset consisting of challenging question-answer-decoy (QAD) sets for long videos (up to 15 minutes).
12 papers · 0 benchmarks
The largest real-world night-time semantic segmentation dataset with pixel-level labels.
12 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
12 papers · 1 benchmark
ONCE-3DLanes is a real-world autonomous driving dataset with lane layout annotation in 3D space.
12 papers · 0 benchmarks
OPA (Object Placement Assessment)
Object-Placement-Assessment (OPA) is a task consisting on verifying whether a composite image is plausible in terms of the object placement.
12 papers · 0 benchmarks
OPUS (open parallel corpus)
OPUS is a growing collection of translated texts from the web.
12 papers · 0 benchmarks
The Overruling dataset is a law dataset corresponding to the task of determining when a sentence is overruling a prior decision.
12 papers · 1 benchmark
P3M-10k (Privacy-Preserving Portrait Matting Dataset)
P3M-10k contains 10421 high-resolution real-world face-blurred portrait images, along with their manually labeled alpha mattes.
12 papers · 1 benchmark
PROST (Physical Reasoning about Objects Through Space and Time)
The PROST (Physical Reasoning about Objects Through Space and Time) dataset contains 18,736 multiple-choice questions made from 14 manually curated templates, covering 10 physical reasoning concepts.
12 papers · 0 benchmarks
ParsiNLU is a comprehensive suite of high-level Natural Language Processing (NLP) tasks for the Persian language.
12 papers · 0 benchmarks
Partial iLIDS is a dataset for occluded person person re-identification.
12 papers · 0 benchmarks
We present PeerQA, a real-world, scientific, document-level Question Answering (QA) dataset.
12 papers · 3 benchmarks
People-Art is an object detection dataset which consists of people in 43 different styles.
12 papers · 2 benchmarks
A dataset that contains 25,017 reading comprehension style examples curated from an existing corpus of 115 website privacy policies.
12 papers · 0 benchmarks
RECON (RECON Outdoor Navigation Dataset)
https://sites.google.com/view/recon-robot/dataset
12 papers · 0 benchmarks
A dataset of color images corrupted by natural noise due to low-light conditions, together with spatially and intensity-aligned low noise images of the same scenes.
12 papers · 1 benchmark
RICE (Remote sensing Image Cloud rEmoving)
RICE is a remote sensing image dataset for cloud removal.
12 papers · 1 benchmark
We managede to collect a real-world rain dataset, named RainDS, includinnumerousus image pairs in various lighting conditions and different scenes.
12 papers · 0 benchmarks
RareAct is a video dataset of unusual actions, including actions like “blend phone”, “cut keyboard” and “microwave shoes”.
12 papers · 1 benchmark
RarePlanes is a unique open-source machine learning dataset from CosmiQ Works and AI.Reverie that incorporates both real and synthetically generated satellite imagery.
12 papers · 0 benchmarks
The Robotic Pushing Dataset is a dataset for video prediction for real-world interactive agents which consists of 59,000 robot interactions involving pushing motions, including a test set with novel objects.
12 papers · 0 benchmarks
SILK (Synth It Like KITTI)
An important factor in advancing autonomous driving systems is simulation.
12 papers · 0 benchmarks
SODA10M is a large-scale object detection benchmark for standardizing the evaluation of different self-supervised and semi-supervised approaches by learning from raw data.
12 papers · 0 benchmarks
SemanticSTF is an adverse-weather point cloud dataset that provides dense point-level annotations and allows to study 3DSS under various adverse weather conditions.
12 papers · 1 benchmark
Semi-iNat is a challenging dataset for semi-supervised classification with a long-tailed distribution of classes, fine-grained categories, and domain shifts between labeled and unlabeled data.
12 papers · 0 benchmarks
SentNoB (SentNoB: A Dataset for Analysing Sentiment on Noisy Bangla Texts)
Social Media User Sentiment Analysis Dataset.
12 papers · 0 benchmarks
This multi-view pant-tilt-zoom-camera (PTZ) dataset features competitive alpine skiers performing giant slalom runs.
12 papers · 1 benchmark
SpaceNet 7 (Multi-Temporal Urban Development SpaceNet Dataset)
Satellite imagery analytics have numerous human development and disaster response applications, particularly when time series methods are involved.
12 papers · 0 benchmarks
An 'in the wild' dataset of 20,580 dog images for which 2D joint and silhouette annotations were collected.
12 papers · 1 benchmark
Synthia-Seq contains 8,000 photo-realistic frames with dense segmentation labels.
12 papers · 0 benchmarks
Co-speech gestures are everywhere.
12 papers · 1 benchmark
TJU-DHD is a high-resolution dataset for object detection and pedestrian detection.
12 papers · 2 benchmarks
The TYO-L (Toyota Light) dataset is part of the Benchmark for 6D Object Pose Estimation (BOP).
12 papers · 0 benchmarks
The Dataset is part of the KELM corpus This is the Wikipedia text--Wikidata KG aligned corpus used to train the data-to-text generation model.
12 papers · 1 benchmark
The TrajNet Challenge represents a large multi-scenario forecasting benchmark.
12 papers · 2 benchmarks
The TrashCan dataset is an instance-segmentation dataset of underwater trash.
12 papers · 0 benchmarks
The UCFRep dataset contains 526 annotated repetitive action videos.
12 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
12 papers · 1 benchmark
We construct the long-tailed version of VOC from its 2012 train-val set.
12 papers · 2 benchmarks
VOT2014 (Visual Object Tracking Challenge 2014)
The dataset comprises 25 short sequences showing various objects in challenging backgrounds.
12 papers · 1 benchmark
ViHSD (Vietnamese Hate Speech Detection Dataset)
This dataset contains 33,400 annotated comments used for hate speech detection on social network sites.
12 papers · 0 benchmarks
VideoMatte240K consists of 484 high-resolution green screen videos and generate a total of 240,709 unique frames of alpha mattes and foregrounds with chroma-key software Adobe After Effects.
12 papers · 0 benchmarks
VideoSet is a large-scale compressed video quality dataset based on just-noticeable-difference (JND) measurement.
12 papers · 0 benchmarks
Consists of over 39,000 images originating from people who are blind that are each paired with five captions.
12 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.