Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 211 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10081–10128 of 12,172
We create a new benchmark called the Smart-City CCTV Violence Detection dataset (SCVD).
1 paper · 0 benchmarks
SD7K (Shadow Document 7K)
SD7K is the only large-scale high-resolution dataset that satisfies all important data features about document shadow currently, which covers a large number of document shadow images.
1 paper · 0 benchmarks
This is a dataset of 306,006 galaxies whose coordinates are taken from the Sloan Digital Sky Survey Data Release 7 and a modified catalogue from Brinchmann+2003 and Wilman+2010.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is associated with the paper published in Scientific Data, titled "SDWPF: A Dataset for Spatial Dynamic Wind Power Forecasting over a Large Turbine Array." You can access the paper:…
1 paper · 0 benchmarks
SE-PEF (Stack Exchange - Personalized Expert Finding)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SE-PQA (SE-PQA: a Resource for Personalized Community Question Answering)
Personalization in Information Retrieval is a topic studied for a long time.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Read more about the dataset here: https://github.com/ServiceNow/seasonal-contrast
1 paper · 0 benchmarks
Click to add a brisef description of the datdaset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A Benchmark Dataset for Deep Learning-based Methods for 3D Topology Optimization.
1 paper · 0 benchmarks
SES (Spanish Emotional Speech)
Currently, an essential point in speech synthesis is the addressing of the variability of human speech.
1 paper · 0 benchmarks
SESIV (SEmantic Salient Instance Video)
SEmantic Salient Instance Video (SESIV) dataset is obtained by augmenting the DAVIS-2017 benchmark dataset by assigning semantic ground-truth for salient instance labels.
1 paper · 0 benchmarks
SF-MASK is a collection made from 20k low-resolution images exported from diverse and heterogeneous datasets, ranging from 7 x 7 to 64 x 64 pixel resolution.
1 paper · 0 benchmarks
Short-Films 20K (SF20K) is the largest publicly available movie dataset.
1 paper · 0 benchmarks
This work contributes a large, complex, and realistic high-quality safety clothing and helmet detection (SFCHD) dataset.
1 paper · 1 benchmark
SFDDD (State Farm Distracted Driver Detection)
We've all been there: a light turns green and the car in front of you doesn't budge.
1 paper · 0 benchmarks
SFIEB (Star_Field_Image_Enhancement_Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The dataset SFU-HW-Objects-v1 contains bounding boxes and object class labels for High Efficiency Video Coding (HEVC) v1 Common Test Conditions (CTC) video sequences.
1 paper · 0 benchmarks
SFU-HW-Tracks is a dataset for Object Tracking on raw video sequences that contains object annotations with unique object identities (IDs) for the High Efficiency Video Coding (HEVC) v1 Common Test Conditions (CTC) sequences.
1 paper · 0 benchmarks
A dataset collected in a set of experiments that involves human participants and a robot.
1 paper · 0 benchmarks
SFpark (San Francisco Park Evaluation)
The San Francisco Municipal Transportation Agency (SFMTA) website provides data collected during the SFpark pilot project.
1 paper · 0 benchmarks
SG-NLG (Schema-Guided Natural Language Generation)
The SG-NLG dataset is a pre-processed version of the DSTC8 Schema-Guided Dialogue SGD dataset, designed specifically for data-to-text Natural Language Generation (NLG).
1 paper · 0 benchmarks
A 3D skeleton-based group activity understanding dataset.
1 paper · 0 benchmarks
Data for Score-Based Generative Models for PET Image Reconstruction.
1 paper · 0 benchmarks
For testing refusal behavior in a cultural setting, we introduce SGXSTest — a set of manually curated prompts designed to measure exaggerated safety within the context of Singaporean culture.
1 paper · 0 benchmarks
SHADR (sythetic SDoH Human Annotated Demographic Robustness dataset (SHADR))
SDoH Human Annotated Demoographic Robustness (SHADR) Dataset Overview The Social determinants of health (SDoH) play a pivotal role in determining patient outcomes.
1 paper · 0 benchmarks
SHAJ (Spoken Hate in the Albanian Jargon)
This is an abusive/offensive language detection dataset for Albanian.
1 paper · 1 benchmark
This dataset is based on the Spiking Heidelberg Digits (SHD) dataset.
1 paper · 1 benchmark
Benchmark for BC Ki-67 stained cell detection and further annotated classification of cells.
1 paper · 0 benchmarks
Finding a correspondence between two shapes is a fundamental task in computer graphics and geometry processing with applications ranging from texture mapping to animation.
1 paper · 0 benchmarks
SHREC'20 (Shape Correspondence with Non-Isometric Deformations)
1 paper · 0 benchmarks
SHRED-ROM (Reduced order modeling with shallow recurrent decoder networks)
SHallow REcurrent Decoder-based Reduced Order Model (SHRED-ROM) is an ultra-hyperreduced order modeling framework aiming at reconstructing high-dimensional data from limited sensor measurements in multiple scenarios.
1 paper · 0 benchmarks
Safety helmet wearing dataset
1 paper · 0 benchmarks
A large collection of human-written natural language questions and their corresponding SPARQL queries over federated bioinformatics knowledge graphs (KGs) collected for several years across different research groups at the SIB Swiss…
1 paper · 0 benchmarks
SICKLE (Satellite Imagery for Cropping annotated with Keyparameter LabEls)
The availability of well-curated datasets has driven the success of Machine Learning (ML) models.
1 paper · 1 benchmark
Smartphone cameras are ubiquitous in daily life, yet their performance can be severely impacted by dirty lenses, leading to degraded image quality.
1 paper · 0 benchmarks
SIDOD is a new, publicly-available image dataset generated by the NVIDIA Deep Learning Data Synthesizer intended for use in object detection, pose estimation, and tracking applications.
1 paper · 0 benchmarks
The Sequence labellIng evaLuatIon benChmark fOr spoken laNguagE (SILICONE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems specifically designed for spoken language.
1 paper · 1 benchmark
SINS is a database of continuous real-life audio recordings in a home environment.
1 paper · 1 benchmark
SIRST-UAVB (Single frame infrared small target dataset - UAV and birds.)
Infrared dim-small target detection has gained increasing importance in both military and civilian applications due to its ability to detect thermal radiation, operate effectively at night, passively sense radiation, and offer strong…
1 paper · 0 benchmarks
SIS (Symmetry inference sentence)
Comprises of 400 naturalistic usages of literature-informed verbs spanning the spectrum of symmetry-asymmetry.
1 paper · 0 benchmarks
We present the SJTU Multispectral Object Detection (SMOD) dataset for detection.
1 paper · 0 benchmarks
This dataset is developed to estimate bearing loads under various operating conditions (rotational speed, axial and radial loads) using data from temperature and vibration sensors.
1 paper · 0 benchmarks
SKINL2 (Light Field Image Dataset of Skin Lesions)
The SKINL2 dataset comprises a total of 376 light fields acquired under similar conditions.
1 paper · 0 benchmarks
I'm writing a description for the SMIC dataset on Papers With Code.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.