Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 213 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10177–10224 of 12,172
STRAT (Spatial TRAnsformation for virtual Try-on)
Spatial TRAnsformation for virtual Try-on (STRAT) dataset contains three subdatasets: STRAT-glasses, STRAT-hat, and STRAT-tie, which correspond to "glasses try-on", "hat try-on", and "tie try-on" respectively.
1 paper · 0 benchmarks
STURM-Flood (STURM-Flood: a curated dataset for deep learning-based flood extent mapping leveraging Sentinel-1 and Sentinel-2 imagery)
The repository hosts the STURM-Flood dataset, an open-access resource designed for flood extent mapping using Sentinel-1 and Sentinel-2 satellite imagery.
1 paper · 0 benchmarks
STVD-PVCD (Partial Video Copy Detection Dataset)
STVD is the largest public dataset on the PVCD task.
1 paper · 1 benchmark
SUDO is a benchmark of 50 real-world malicious tasks designed to evaluate LLM-based computer agents in live desktop and web environments.
1 paper · 1 benchmark
SUDOER (System/User Dataset for Obedience Evaluation in Responses)
The dataset aims to provide system prompts and user prompts for assistant.
1 paper · 0 benchmarks
A RGB-D dataset converted from SUN-RGBD into COCO-style instance segmentation format.
1 paper · 2 benchmarks
SURL (IMC 2020 Curlie URL Dataset)
"Identifying Sensitive URLs at Web-Scale" dataset at IMC20
1 paper · 0 benchmarks
Added information about the subject's body height and volumes of 14 individual body parts.
1 paper · 0 benchmarks
SUT (SUT: a new multi-purpose synthetic dataset for Farsi document image analysis)
This paper introduces a new large-scale dataset for Farsi document images, named SUT, which aims to tackle the challenges associated with obtaining diverse and substantial ground-truth data for supervised models in document image analysis…
1 paper · 2 benchmarks
A labeled dataset that presents fake news surrounding the conflict in Syria.
1 paper · 0 benchmarks
SVLD (Social Vision and Language Dataset)
The social vision and language dataset is a large-scale multimodal dataset designed for research into social contextual learning.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SVRT (Synthetic Visual Reasoning Task)
The Synthetic Visual Reasoning Test (SVRT) is a series of 23 classification problems involving images of randomly generated shapes.
1 paper · 0 benchmarks
SWS (Smart Word Suggestions Benchmark)
Smart Word Suggestions (SWS) is a task and benchmark.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SYNTHIA-PANO is the panoramic version of SYNTHIA dataset.
1 paper · 0 benchmarks
The SYSU-CEUS dataset consists of three types of Focal liver lesions (FLLs): 186 HCC instances, 109 HEM instances and 58 FNH instances (i.e.,186 malignant instances and 167 benign instances).
1 paper · 0 benchmarks
SYSU-MM01-C is an evaluation set that consists of algorithmically generated corruptions applied to the SYSU-MM01 test-set, and especially to both the visible and the thermal data.
1 paper · 0 benchmarks
S_B_D (Synthetic Barcode Dataset)
100,000 LR synthetic barcode datasets along with their corresponding bounding boxes ground truth masks.
1 paper · 0 benchmarks
SaGA (The Bielefeld Speech and Gesture Alignment Corpus (SaGA))
The primary data of the SaGA corpus are made up of 25 dialogs of interlocutors (50), who engage in a spatial communication task combining direction-giving and sight description.
1 paper · 0 benchmarks
SaRNet is a single class dataset consisting of tiles of satellite imagery labeled with potential 'targets'.
1 paper · 0 benchmarks
This dataset provides harmful queries and their corresponding safe and harmful responses.
1 paper · 0 benchmarks
SafeEdit encompasses 4,050 training, 2,700 validation, and 1,350 test instances.
1 paper · 0 benchmarks
Speech Recognition Dataset for Oromo Language.
1 paper · 1 benchmark
A dataset for grounded language learning that consists of navigational instructions and actions in a maze-like environment.
1 paper · 0 benchmarks
This dataset contains nine video sequences captured by a webcam for salient closed boundary tracking evaluation.
1 paper · 0 benchmarks
Salient-KITTI is a saliency map prediction dataset based on KITTI.
1 paper · 0 benchmarks
Dataset contains cumulative reported cases, hospital admission and discharge, and mortality data as parsed from the publicly available press releases by the Ministry of Health and National Emergency Operations Centre (NEOC) of the…
1 paper · 0 benchmarks
We have uploaded a sample dataset for training and testing Back-Projection Diffusion.
1 paper · 0 benchmarks
The sample EEG dataset consists of the newborn EEG data recorded for the work published as: Buiatti M.
1 paper · 0 benchmarks
Sandhi Kosh is the first Sanskrit Sandhi Benchmark created to evaluate the Sanskrit Sandhi Tools.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SatBird (a Dataset for Bird Species Distribution Modeling using Remote Sensing and Citizen Science Data)
SatBird is a dataset and benchmark for the task of predicting bird species encounter rates jointly at a specific location using remote sensing data.
1 paper · 0 benchmarks
SatIQ Model Weights (Model Weights for "Watch This Space: Securing Satellite Communication through Resilient Transmitter Fingerprinting")
Model weights for use with the SatIQ fingerprinting models used in the paper “Watch This Space: Securing Satellite Communication through Resilient Transmitter Fingerprinting”.
1 paper · 0 benchmarks
The Satellite dataset forms a practical VFL scenario for location identification based on satellite imagery.
1 paper · 0 benchmarks
The satire dataset is a new multi-modal dataset of satirical and regular news articles.
1 paper · 0 benchmarks
Scan Entities in 3D (ScanEnts3D) is a large-scale dataset which provides explicit correspondences between 369k objects across 84k natural referentural sentences, covering 705 real-world scenes.
1 paper · 0 benchmarks
SceneNet-RGBD is a synthetic dataset containing large-scale photorealistic renderings of indoor scene trajectories with pixel-level annotations.
1 paper · 0 benchmarks
The “Mental Health” forum was used, a forum dedicated to people suffering from schizophrenia and different mental disorders.
1 paper · 1 benchmark
SciCo (Scientific Concept Induction Corpus)
SciCo is an expert-annotated dataset for hierarchical CDCR (cross-document coreference resolution) for concepts in scientific papers, with the goal of jointly inferring coreference clusters and hierarchy between them.
1 paper · 0 benchmarks
ScienceExamCER is a collection of resources for studying explanation-centered inference, including explanation graphs for 1,680 questions, with 4,950 tablestore rows, and other analyses of the knowledge required to answer elementary and…
1 paper · 0 benchmarks
This resource contains 10.5 million paragraphs with associated statement labels, realized as one paragraph per file, one sentence per line.
1 paper · 0 benchmarks
A collection of long-running (80+ episodes) science fiction TV show synopses, scraped from Fandom.com wikis.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Scroll Readability Dataset contains scroll interactions of 598 participants reading advanced and elementary texts from the OneStopEnglish corpus.
1 paper · 0 benchmarks
Search4Code is a large-scale web query based dataset of code search queries for C# and Java.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.