Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 145 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6913–6960 of 12,172
https://huggingface.co/papers/2502.20730
2 papers · 0 benchmarks
The first dataset contains annotated natural language queries (i.e.
2 papers · 0 benchmarks
Spades (Semantic PArsing of DEclarative Sentences)
Datasets Spades contains 93,319 questions derived from clueweb09 sentences.
2 papers · 0 benchmarks
Sparrow (Sparrow-V0: A Reinforcement Learning Friendly Simulator for Mobile Robot)
Sparrow-V0: A Reinforcement Learning Friendly Simulator for Mobile Robot Features: Vectorizable (Enable fast data collection; Single environment is also supported) Domain Randomization (control interval, control delay, maximum velocity,…
2 papers · 0 benchmarks
Spatial LibriSpeech is spatial audio dataset with over 650 hours of 19-channel audio, first-order ambisonics, and optional distractor noise.
2 papers · 0 benchmarks
SpatialSense Benchmark is a dataset specializing in spatial relation recognition which captures a broad spectrum of such challenges, allowing for proper benchmarking of computer vision techniques.
2 papers · 0 benchmarks
The speech accent archive uniformly presents a large set of speech samples from a variety of language backgrounds.
2 papers · 1 benchmark
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
This resource is designed to allow for research into Natural Language Generation.
2 papers · 0 benchmarks
The Stack Exchange dataset is a collection of data from various Stack Exchange sites, including Stack Overflow, Mathematics, Super User, and many others.
2 papers · 1 benchmark
The dataset has two years of user awards on a question-answering website: each user received a sequence of badges and there are 22 different kinds of badges in total.
2 papers · 1 benchmark
The "Stance Detection in COVID-19 Tweets" dataset represents an evolution of stance detection research, tailored to address the unique and urgent challenges presented by the COVID-19 pandemic.
2 papers · 0 benchmarks
Electrophysiological data from implanted electrodes in the human brain are rare, and therefore scientific access to it has remained somewhat exclusive.
2 papers · 1 benchmark
Schema2QA is the first large question answering dataset over real-world Schema.org data.
2 papers · 0 benchmarks
Stanford-ECM is an egocentric multimodal dataset which comprises about 27 hours of egocentric video augmented with heart rate and acceleration data.
2 papers · 0 benchmarks
The Stanford 40 Action Dataset contains images of humans performing 40 actions.
2 papers · 1 benchmark
StarData is a StarCraft: Brood War replay dataset, with 65,646 games.
2 papers · 0 benchmarks
StereoMSI comprises of 350 registered colour-spectral image pairs.
2 papers · 0 benchmarks
A multimodal empathetic dialogue dataset.
2 papers · 0 benchmarks
StoryBench (StoryBench: A Multifaceted Benchmark for Continuous Story Visualization)
StoryBench is a multi-task benchmark to reliably evaluate the ability of text-to-video models to generate stories from a sequence of captions and their duration.
2 papers · 1 benchmark
StreaksYoloDataset, is a set of raw astronomical images captured with smart telescopes and annotated with the positions of streaks that are effectively in the images.
2 papers · 0 benchmarks
Text-Vison Cross-Modal Place Recognition Dataset
2 papers · 0 benchmarks
StreetTryOn, the new in-the-wild Virtual Try-On dataset, consists of 12,364 and 2,089 street person images for training and validation, respectively.
2 papers · 1 benchmark
Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study File Descriptions File | Description --- | --- commitcategorizations.csv | Categorizations for the commits in our dataset.
2 papers · 0 benchmarks
StyleKQC is a style-variant paraphrase corpus for korean questions and commands.
2 papers · 0 benchmarks
SuHiFiMask (Surveillance High-Fidelity Mask) extends FAS to real surveillance scenes rather than mimicking low-resolution images and surveillance environments.
2 papers · 0 benchmarks
SuMe (A Dataset Towards Summarizing Biomedical Mechanisms)
Can language models read biomedical texts and explain the biomedical mechanisms discussed?
2 papers · 0 benchmarks
SubEdits is a human-annnoated post-editing dataset of neural machine translation outputs, compiled from in-house NMT outputs and human post-edits of subtitles form Rakuten Viki.
2 papers · 0 benchmarks
This is a discourse dataset with multiple and subjective interpretations of English conversation in the form of perceived conversation acts and intents.
2 papers · 0 benchmarks
The dataset represents data generated from a commonly used model in population genetics.
2 papers · 0 benchmarks
Introduction This dataset supports Ye et al.
2 papers · 0 benchmarks
SuperCaustics is a simulation tool made in Unreal Engine for generating massive computer vision datasets that include transparent objects.
2 papers · 0 benchmarks
This repository contains processed data and result files for the paper "Revealing drivers and risks for power grid frequency stability with explainable AI".
2 papers · 0 benchmarks
The training subset consists of 15 robotic nephrectomy procedures captured on the da Vinci X or Xi system.
2 papers · 0 benchmarks
The SweRec dataset in ScandEval is a Swedish language dataset used for text classification tasks.
2 papers · 0 benchmarks
Uses a platform with 77 candies and sweets to rank.
2 papers · 0 benchmarks
A Stated Preference Survey on mode choice https://transp-or.epfl.ch/documents/technicalReports/CSSwissmetroDescription.pdf
2 papers · 0 benchmarks
SymbolicData (A Tree-based Symbolic Dataset For Symbolic Regression)
This dataset is a collection of input-label pairs where each input is in the form of a numerical dataset, itself a set of input and output pairs {(x, y)}, and the corresponding label is a string encoding the symbolic expression governing…
2 papers · 0 benchmarks
SynMirror consists of samples rendered from 3D assets of two widely used 3D object datasets - Objaverse and Amazon Berkeley Objects (ABO) placed in front of a mirror in a virtual blender environment.
2 papers · 0 benchmarks
SynthDerm is a synthetically generated dataset inspired by the real-world characteristics of melanoma skin lesions in dermatology settings.
2 papers · 0 benchmarks
Event cameras are sensors that are inspired by biological systems and specialize in capturing changes in brightness.
2 papers · 1 benchmark
The SynthSOD dataset contains more than 47 hours of multitrack music obtained by synthesizing orchestra and ensemble pieces from the Symbolic Orchestral Database (SOD) using Spitfire BBC Symphony Orchestra Professional Library.
2 papers · 0 benchmarks
Synthetic OD data to mimic data showed in the application of the paper.
2 papers · 1 benchmark
TAP (Traffic Accident Prediction data repository)
The Traffic Accident Prediction (TAP) data repository offers extensive coverage for 1,000 US cities (TAP-city) and 49 states (TAP-state), providing real-world road structure data that can be easily used for graph-based machine learning…
2 papers · 0 benchmarks
TAPVid-3D is a dataset and benchmark for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D).
2 papers · 0 benchmarks
TAS-NIR is a VIS+NIR dataset of semantically annotated images in unstructured outdoor environments.
2 papers · 0 benchmarks
The TAU-NIGENS Spatial Sound Events 2020 dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and…
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.