Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 105 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4993–5040 of 12,172
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
Machine-learning Data Set Prepared from NASA Solar Dynamics Observatory Mission data.
4 papers · 0 benchmarks
SDSD-outdoor (Seeing Dynamic Scene in the Dark: High-Quality Video Dataset with Mechatronic Alignment)
Seeing Dynamic Scene in the Dark: High-Quality Video Dataset with Mechatronic Alignment
4 papers · 1 benchmark
SLING consists of 38K minimal sentence pairs in Mandarin Chinese grouped into 9 high-level linguistic phenomena.
4 papers · 0 benchmarks
SOD4SB (Small Object Detection for Spotting Birds)
The Small Object Detection for Spotting Birds (SOD4SB) dataset is a dataset consisting of 39,070 images including 137,121 bird instances.
4 papers · 2 benchmarks
A dataset for urban sound tagging with spatiotemporal information.
4 papers · 0 benchmarks
SPACE is a large-scale opinion summarization benchmark for the evaluation of unsupervised summarizers.
4 papers · 1 benchmark
STAIR Captions is a large-scale dataset containing 820,310 Japanese captions.
4 papers · 0 benchmarks
A novel traffic flow dataset from publicly available web cameras in the suburbs of Chicago, IL.
4 papers · 0 benchmarks
This is a classification problem to distinguish between a signal process which produces supersymmetric particles and a background process which does not.
4 papers · 0 benchmarks
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
A new large-scale dataset along with an open-source library for SVG manipulation.
4 papers · 0 benchmarks
SWAT A7 (Secure Water Treatment (SWaT))
11 days of continuous operation: 7 under normal operation and 4 days with attack scenarios: + Collected network traffic & all the values obtained from all the 51 sensors and actuators + Data labelled according to normal and abnormal…
4 papers · 0 benchmarks
SWINySEG (Singapore Whole sky Nychthemeron Image SEGmentation Database)
The SWINySEG dataset contains 6768 daytime- and nighttime-images of sky/cloud patches along with their corresponding binary ground truth maps.
4 papers · 1 benchmark
SWORD ('Scenes with occluded regions' dataset)
The new dataset contains around 1,500 train videos and 290 test videos, with 50 frames per video on average.
4 papers · 1 benchmark
SYMON (Synopses of Movie Narratives)
Contains 5,193 video summaries of popular movies and TV series.
4 papers · 0 benchmarks
SYSU-MM01-C is an evaluation set that consists of algorithmically generated corruptions applied to the SYSU-MM01 test-set.
4 papers · 1 benchmark
Taxi speed data in 15min interval from 156 sensors on major roads of Luohu District in Shenzhen, China, from Jan.
4 papers · 1 benchmark
Physics-based simulated garments on top of SMPL bodies.
4 papers · 0 benchmarks
The ScaLA dataset is a linguistic acceptability dataset for the Scandinavian languages, including Danish, Norwegian Bokmål, Norwegian Nynorsk, Swedish, Icelandic, and Faroese.
4 papers · 0 benchmarks
ScenicOrNot (SoN) is a dataset of 185,548 images with associated natural beauty rating histograms.
4 papers · 0 benchmarks
SciDuet is a dataset for training and benchmarking models for automating document-to-slides generation.
4 papers · 0 benchmarks
SecQA is a specialized dataset created for the evaluation of Large Language Models (LLMs) in the domain of computer security.
4 papers · 0 benchmarks
Symlink is a SemEval shared task of extracting mathematical symbols and their descriptions from LaTeX source of scientific documents.
4 papers · 1 benchmark
Sentiment140 is a dataset that allows you to discover the sentiment of a brand, product, or topic on Twitter.
4 papers · 1 benchmark
Separated COCO is automatically generated subsets of COCO val dataset, collecting separated objects for a large variety of categories in real images in a scalable manner, where target object segmentation mask is separated into distinct…
4 papers · 1 benchmark
The expansion of social networks has accelerated the transmission of information and news at every communities.
4 papers · 1 benchmark
ShapenetRenderer is an extension of the ShapeNet Core dataset which has more variation in camera angles.
4 papers · 0 benchmarks
A dataset of real distributional shift across multiple large-scale tasks.
4 papers · 0 benchmarks
SoftAttributes (SoftAttributes: Relative movie attribute dataset for soft attributes)
The dataset consists of sets of movie titles, with each set annotated with a single English soft attribute (subjective descriptive property, such as 'confusing' or 'romantic') and a reference movie.
4 papers · 0 benchmarks
Solar-Power (Solar Power Data for Integration Studies (Alabama))
Solar Power Data for Integration Studies NREL's Solar Power Data for Integration Studies are synthetic solar photovoltaic (PV) power plant data points for the United States representing the year 2006.
4 papers · 2 benchmarks
Some Like it Hoax is a fake news detection dataset consisting of 15,500 Facebook posts and 909,236 users.
4 papers · 0 benchmarks
Something-Something-100 is a dataset split created from Something-Something V2.
4 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
SpeechInstruct is a large-scale cross-modal speech instruction dataset.
4 papers · 0 benchmarks
SportsPose (SportsPose - A Dynamic 3D sports pose dataset)
Accurate 3D human pose estimation is essential for sports analytics, coaching, and injury prevention.
4 papers · 0 benchmarks
A growing number of papers are published in the area of superconducting materials science.
4 papers · 1 benchmark
Swiss3DCities is a dataset that is manually annotated for semantic segmentation with per-point labels, and is built using photogrammetry from images acquired by multirotors equipped with high-resolution cameras.
4 papers · 0 benchmarks
This dataset contains a variety of common urban road objects scanned with a Velodyne HDL-64E LIDAR, collected in the CBD of Sydney, Australia.
4 papers · 1 benchmark
SynthPAI (SynthPAI: A Synthetic Dataset for Personal Attribute Inference)
SynthPAI was created to provide a dataset that can be used to investigate the personal attribute inference (PAI) capabilities of LLM on online texts.
4 papers · 1 benchmark
Synthinel-1 is a collection of synthetic overhead imagery with full pixel-wise building segmentation labels.
4 papers · 0 benchmarks
The Szeged Treebank is the largest fully manually annotated treebank of the Hungarian language.
4 papers · 0 benchmarks
TAC 2010 is a dataset for summarization that consists of 44 topics, each of which is associated with a set of 10 documents.
4 papers · 1 benchmark
TAU Spatial Sound Events 2019 consists of 2 datasets: Ambisonic (FOA) and Microphone Array (MIC), of identical sound scenes with the only difference in the format of the audio.
4 papers · 0 benchmarks
The TAU-NIGENS Spatial Sound Events 2021 dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and…
4 papers · 1 benchmark
A morpho-syntactically annotated Tunisian Arabish Corpus (TArC).
4 papers · 0 benchmarks
The largest and most realistic dataset available for TCC.
4 papers · 0 benchmarks
TCG (Traffic Control Gesture)
The TCG dataset is used to evaluate Traffic Control Gesture recognition for autonomous driving.
4 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.