Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 62 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2929–2976 of 12,172
Ohsumed includes medical abstracts from the MeSH categories of the year 1991.
11 papers · 2 benchmarks
Open PI is the first dataset for tracking state changes in procedural text from arbitrary domains by using an unrestricted (open) vocabulary.
11 papers · 0 benchmarks
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
11 papers · 2 benchmarks
P-DukeMTMC-reID is a modified version based on DukeMTMC-reID dataset.
11 papers · 1 benchmark
PATS (Pose Audio Transcript Style)
PATS dataset consists of a diverse and large amount of aligned pose, audio and transcripts.
11 papers · 0 benchmarks
PTR is a new large-scale diagnostic visual reasoning dataset for research around part-based conceptual, relational and physical reasoning.
11 papers · 0 benchmarks
ParCorFull (Parallel Corpus Annotated with Full Coreference)
ParCorFull is a parallel corpus annotated with full coreference chains that has been created to address an important problem that machine translation and other multilingual natural language processing (NLP) technologies face -- translation…
11 papers · 0 benchmarks
Perspectrum is a dataset of claims, perspectives and evidence, making use of online debate websites to create the initial data collection, and augmenting it using search engines in order to expand and diversify the dataset.
11 papers · 1 benchmark
A large-scale collection of visually-grounded, task-oriented dialogues in English designed to investigate shared dialogue history accumulating during conversation.
11 papers · 0 benchmarks
Place Pulse is a crowdsourcing effort that aims to map which areas of a city are perceived as safer, livelier, wealthier, more active, beautiful and friendly.
11 papers · 1 benchmark
ProtoQA is a question answering dataset for training and evaluating common sense reasoning capabilities of artificial intelligence systems in such prototypical situations.
11 papers · 0 benchmarks
RGB-D-D is a large-scale dataset for depth map super-resolution (SR).
11 papers · 0 benchmarks
RMFD (Real-World Masked Face Dataset)
Real-World Masked Face Dataset (RMFD) is a large dataset for masked face detection.
11 papers · 0 benchmarks
A large-scale non-homogeneous remote sensing image dehazing dataset
11 papers · 1 benchmark
RailSem19 (RailSem19: A Dataset for Semantic Rail Scene Understanding)
RailSem19 offers 8500 unique images taken from a the ego-perspective of a rail vehicle (trains and trams).
11 papers · 0 benchmarks
ReCAM (SemEval-2021 Task 4: Reading Comprehension of Abstract Meaning)
Tasks Our shared task has three subtasks.
11 papers · 1 benchmark
A large-scale dataset of ~29.5K rain/rain-free image pairs that covers a wide range of natural rain scenes.
11 papers · 0 benchmarks
The RoboTAP dataset follows the same annotation format as TAP-Vid, but is released as an addition to TAP-Vid.
11 papers · 0 benchmarks
A collection that allows researchers to approach the extremely challenging problem of description generation using relatively simple non-parametric methods and produces surprisingly effective results.
11 papers · 0 benchmarks
The SD-198 dataset contains 198 different diseases from different types of eczema, acne and various cancerous conditions.
11 papers · 0 benchmarks
SEDE (Stack Exchange Data Explorer)
SEDE is a dataset comprised of 12,023 complex and diverse SQL queries and their natural language titles and descriptions, written by real users of the Stack Exchange Data Explorer out of a natural interaction.
11 papers · 1 benchmark
SHHS (Sleep Heart Health Study)
The Sleep Heart Health Study (SHHS) is a multi-center cohort study implemented by the National Heart Lung & Blood Institute to determine the cardiovascular and other consequences of sleep-disordered breathing.
11 papers · 1 benchmark
SLOPER4D is a novel scene-aware dataset collected in large urban environments to facilitate the research of global human pose estimation (GHPE) with human-scene interaction in the wild.
11 papers · 1 benchmark
SMAC-Exp (StarCraft Multi-Agent Exploration Challenge)
The StarCraft Multi-Agent Challenges+ requires agents to learn completion of multi-stage tasks and usage of environmental factors without precise reward functions.
11 papers · 2 benchmarks
SOMOS (The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis)
The SOMOS dataset is a large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples.
11 papers · 0 benchmarks
So2Sat LCZ42 consists of local climate zone (LCZ) labels of about half a million Sentinel-1 and Sentinel-2 image patches in 42 urban agglomerations (plus 10 additional smaller areas) across the globe.
11 papers · 1 benchmark
StylePTB is a fine-grained text style transfer benchmark.
11 papers · 0 benchmarks
Synbols is a dataset generator designed for probing the behavior of learning algorithms.
11 papers · 0 benchmarks
TACO is a growing image dataset of waste in the wild.
11 papers · 0 benchmarks
The first large demoire dataset.
11 papers · 1 benchmark
Existing benchmarks for temporal QA focus on a single information source (either a KB or a text corpus), and include only few questions with implicit constraints.
11 papers · 1 benchmark
TRIPOD (TuRnIng POint Dataset)
TRIPOD contains screenplays and plot synopses with turning point (TP) annotations for 99 movies.
11 papers · 0 benchmarks
Overall duration per microphone: about 36 hours (31 hrs train / 2.5 hrs dev / 2.5 hrs test) Count of microphones: 3 (Microsoft Kinect, Yamaha, Samson) Count of wave-files per microphone: about 14500 Overall count of participations: 180…
11 papers · 1 benchmark
TUM monoVO is a dataset for evaluating the tracking accuracy of monocular Visual Odometry (VO) and SLAM methods.
11 papers · 0 benchmarks
TaPaCo is a freely available paraphrase corpus for 73 languages extracted from the Tatoeba database.
11 papers · 0 benchmarks
Talk The Walk is a large-scale dialogue dataset grounded in action and perception.
11 papers · 0 benchmarks
TaxiNLI is a dataset collected based on the principles and categorizations of the aforementioned taxonomy.
11 papers · 0 benchmarks
TimeDial presents a crowdsourced English challenge set, for temporal commonsense reasoning, formulated as a multiple choice cloze task with around 1.5k carefully curated dialogs.
11 papers · 0 benchmarks
The ToughTables (2T) dataset was created for the SemTab challenge and includes 180 tables in total.
11 papers · 4 benchmarks
UDIS-D (Unsupervised Deep Image Stitching Dataset)
UDIS-D is a large image dataset for image stitching or image registration.
11 papers · 0 benchmarks
UnrealEgo is a dataset that provides in-the-wild stereo images with a large variety of motions for 3D human pose estimation.
11 papers · 1 benchmark
Over the past few years a number of research groups have made rapid advances in remote PPG methods for estimating heart rate from digital video and obtained impressive results.
11 papers · 0 benchmarks
VLEP (Video-and-Language Event Prediction)
VLEP contains 28,726 future event prediction examples (along with their rationales) from 10,234 diverse TV Show and YouTube Lifestyle Vlog video clips.
11 papers · 1 benchmark
VNBench is a comprehensive benchmark suite for video generative models, which evaluates video generation quality across specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods.
11 papers · 1 benchmark
VOST consists of more than 700 high-resolution videos, captured in diverse environments, which are 20 seconds long on average and densely labeled with instance masks.
11 papers · 0 benchmarks
The Video2GIF dataset contains over 100,000 pairs of GIFs and their source videos.
11 papers · 0 benchmarks
VinDr-CXR is an open large-scale dataset of chest X-rays with radiologist’s annotations.
11 papers · 0 benchmarks
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
11 papers · 9 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.