Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 214 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10225–10272 of 12,172
This data set includes all raw data (e.g., collected certificates) of the WWW 2021 paper "Security of Alerting Authorities in the WWW: Measuring Namespaces, DNSSEC, and Web PKI".
1 paper · 0 benchmarks
SegSub (SegSub: Evaluating Robustness to Knowledge Conflicts and Hallucinations in Vision-Language Models)
This research introduces \segsub, a framework for applying targeted image perturbations to investigate VLM resilience against knowledge conflicts.
1 paper · 0 benchmarks
This dataset is obtained during an ICON project (2017-2018) in collaboration with KU Leuven (ESAT-STADIUS), UZ Leuven, UCB, Byteflies and Pilipili.
1 paper · 0 benchmarks
This dataset is the image stimulus pool of 50 deepfake and 50 real images, used for the experiment in the study titled "Testing Human Ability To Detect Deepfake Images of Human Faces".
1 paper · 0 benchmarks
Please refer to the Zenodo page for a detailed description: https://zenodo.org/records/15665101
1 paper · 0 benchmarks
Autism Spectrum Disorders (ASD), often referred to as autism, are neurological disorders characterised by deficits in cognitive skills, social and communicative behaviours.
1 paper · 1 benchmark
SemEval-2016 Task 6, titled "Stance Detection in Tweets," provides a specialized dataset for the computational linguistics and natural language processing (NLP) communities to explore and analyze users' positions towards certain targets,…
1 paper · 0 benchmarks
Dataset Card for SemTabNet This dataset accompanies the following paper: Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar…
1 paper · 1 benchmark
Test dataset for Semantic Segmentation.
1 paper · 0 benchmarks
This dataset contains 547 social media images taken in the aftermath of various earthquakes.
1 paper · 0 benchmarks
SemanticSugarBeets, a novel and high-quality dataset containing 953 monocular RGB images and 2920 annotations of sugar beets, enables a wide range of learning tasks including object detection, semantic segmentation, instance segmentation…
1 paper · 0 benchmarks
SemanticUSL is a dataset for domain adaptation for LiDAR point cloud semantic segmentation.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SEN2VENµS is an open dataset for the super-resolution of Sentinel-2 images by leveraging simultaneous acquisitions with the VENµS satellite.
1 paper · 2 benchmarks
Sen4AgriNet (A Sentinel-2 multi-year, multi-country benchmark dataset for crop classification and segmentation with deep learning)
A Sentinel-2 based time series multi country benchmark dataset, tailored for agricultural monitoring applications with Machine and Deep Learning.
1 paper · 0 benchmarks
SensoDat is a dataset of self-driving car simulation data (30K executed simulations).
1 paper · 0 benchmarks
The dataset is based on a debate.org crawl.
1 paper · 0 benchmarks
This is a dataset for 3-way sentiment classification of reviews (negative, neutral, positive).
1 paper · 1 benchmark
SentimentArcs’ reference corpus for novels consists of 25 narratives selected to create a diverse set of well recognized novels that can serve as a benchmark for future studies.
1 paper · 0 benchmarks
This dataset includes 2.133.324 reflectance water spectra which were manually extracted by visual observation from 30 Sentinel 2 level 1C satellite images.
1 paper · 0 benchmarks
Trip duration is the most fundamental measure in all modes of transportation.
1 paper · 0 benchmarks
Sequence Consistency Evaluation (SCE) consists of a benchmark task for sequence consistency evaluation (SCE).
1 paper · 0 benchmarks
This is the sequential instructions dataset from Understanding the Effects of RLHF on LLM Generalisation and Diversity.
1 paper · 0 benchmarks
This dataset accompanies the linked SerialTrack paper and provides test case data (2D/3D, varying particle density) across a range of synthetic and experimental imaging modalities.
1 paper · 0 benchmarks
Large collection of accumulated shadow tiles for over 100 cities (in 6 continents), computed using Deep Umbra.
1 paper · 0 benchmarks
ShadowLink dataset is designed to evaluate the impact of entity overshadowing on the task of entity disambiguation.
1 paper · 0 benchmarks
The ShapeIt dataset introduced by Alper et al.
1 paper · 0 benchmarks
The synthetic ShapeNet intrinsic image decomposition dataset of 90,000 images.
1 paper · 0 benchmarks
The ShapeNet-Skeleton dataset has ground-truth skeleton point sets and skeletal volumes for object instances in the ShapeNet dataset.
1 paper · 0 benchmarks
The problem here is to predict whether a share price will show an exceptional rise after quarterly announcement of the Earning Per Share based on the price movement of that share price on the proceeding 60 days?
1 paper · 0 benchmarks
This repository contains documentation for the dataset that accompanies our ICPE 2025 paper, "Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads".
1 paper · 0 benchmarks
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts 🔥 Key Features - 3000+ hours of synthetic speech - Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including: - 📖 Reading Style - 🎙️…
1 paper · 0 benchmarks
To construct such a dataset, a straightforward approach was scraping images from the web.
1 paper · 1 benchmark
ShopTC-100K Dataset The ShopTC-100K dataset is collected using TermMiner, an open-source data collection and topic modeling pipeline introduced in the paper: Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable…
1 paper · 0 benchmarks
Short BBC Pose contains five one-hour-long videos with sign language signers each with different sleeve length (in contrast to the BBC pose and Extended BBC Pose, which only contain signers with moderately long sleeves).
1 paper · 0 benchmarks
In this Adjudicator ScoresShort Stories and Written Reflections folder: Four files from four student participants of the contest.
1 paper · 0 benchmarks
The proposed dataset includes 1,309 short text instances from Adobe Spark.
1 paper · 0 benchmarks
ShortPersianEmo is a new data set for emotion recognition in Persian short texts.
1 paper · 1 benchmark
ShuttleSet22 is a badminton singles dataset which is collected from high-ranking matches in 2022.
1 paper · 0 benchmarks
SidechainNet is a protein structure prediction dataset that directly extends ProteinNet.
1 paper · 0 benchmarks
The database consists of EEG recordings of 14 patients acquired at the Unit of Neurology and Neurophysiology of the University of Siena.
1 paper · 0 benchmarks
Procedural videos show step-by-step demonstrations of tasks like recipe preparation.
1 paper · 0 benchmarks
Facial electromyography recordings during both silent and vocalized speech.
1 paper · 0 benchmarks
The SimBEV dataset is a collection of 320 scenes spread across all 11 CARLA maps and contains data from a variety of sensors, including five camera types (RGB, semantic segmentation, instance segmentation, depth, and optical flow), lidar,…
1 paper · 3 benchmarks
SimGas (Computer Simulated Gas Leakage Segmentation)
This dataset consists of computer-generated images for gas leakage segmentation.
1 paper · 2 benchmarks
SimpEvalASSET is a dataset for learning learnable metrics using modern language models.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.