Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 212 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10129–10176 of 12,172
The following files comprise 19 sets of 40,000 images, each set corresponding to a different rendering sigma as described in the paper.
1 paper · 0 benchmarks
SMOT (Single sequence-Multi Objects Training)
The SMOT dataset, Single sequence-Multi Objects Training, is collected to represent a practical scenario of collecting training images of new objects in the real world, i.e.
1 paper · 0 benchmarks
This paper introduces CPIR-MR (Chained Prompting for Improved Readability of Medical Reports), a method designed to simplify complex chest X-ray reports for better patient understanding.
1 paper · 0 benchmarks
SOEVAL is created by us by mining questions from StackOverflow.
1 paper · 0 benchmarks
The SOLO Corpus comprises over 4 million English tweets, each of which contains at least one of the following tokens: solitude, lonely, and loneliness.
1 paper · 0 benchmarks
SOMD (SOftware Mention Detection)
The dataset contains the training and test data for the SOftware Mention Detection challenge.
1 paper · 0 benchmarks
SOMPT22 (Surveillance Oriented Multi-Pedestrian Tracking Dataset (SOMPT22))
SOMPT22 is a multi-object tracking (MOT) benchmark focused on surveillance-style pedestrian tracking.
1 paper · 0 benchmarks
SOTIF-PCOD is a dataset generated using the CARLA simulator, specifically designed for Safety of the Intended Functionality (SOTIF) research.
1 paper · 0 benchmarks
Curated QA Benchmark on State of the Union Address 2023.
1 paper · 0 benchmarks
SOTVerse is a user-defined task space of single object tracking.
1 paper · 0 benchmarks
Dataset for Land Cover segmentation from sparse labels, using Sentinel-2 as source imagery.
1 paper · 0 benchmarks
Contains three types of 2D-3D reasoning tasks on view consistency, camera pose, and shape generation, with increasing difficulty.
1 paper · 0 benchmarks
SPARKESX (Single-dish PARKES for finding the uneXpected)
We present the Single-dish PARKES data sets for finding the uneXpected (SPARKESX), a compilation of real and simulated high-time resolution observations.
1 paper · 0 benchmarks
SPAVE-28G (Signal Propagation Analyses in V2X Ecosystems (S.P.A.V.E) at 28 GHz on the NSF POWDER testbed)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This paper details the design of an autonomous alignment and tracking platform to mechanically steer directional horn antennas in a sliding correlator channel sounder setup for 28 GHz V2X propagation modeling.
1 paper · 0 benchmarks
The dataset contains both RGB and depth images, and the data from two accelerometers, together with ground truth calorie values from a calorimeter for calorie expenditure estimation in home environments.
1 paper · 0 benchmarks
The SPI dataset consists of force-controlled industrial robot data for training shadow program inversion (SPI) models.
1 paper · 0 benchmarks
SPIQA Dataset Card Dataset Details Dataset Name: SPIQA (Scientific Paper Image Question Answering) Paper: SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers Github: SPIQA eval and metrics code repo Dataset Summary:…
1 paper · 0 benchmarks
SPKL (Seasonal Parking Lot Dataset)
The SPKL dataset contains 1203 images of parking lots divided into 11 categories regarding vision conditions (including the 'winter' category absent in other datasets at the time of publishing).
1 paper · 1 benchmark
SPOT-10 (Animal Pattern Benchmark Dataset for Machine Learning Algorithms)
The SPOTS-10 dataset is an extensive collection of grayscale images showcasing diverse patterns found in ten animal species.
1 paper · 1 benchmark
SPRIGHT is the first, large-scale vision-language dataset that focuses on spatial relationships.
1 paper · 0 benchmarks
This is a cleaned version of the dataset introduced in Kaggle by user SAJID576.
1 paper · 0 benchmarks
Stanford Question Answering Dataset (SQuAD) into Spanish.
1 paper · 0 benchmarks
SQuAD-it is derived from the SQuAD dataset and it is obtained through semi-automatic translation of the SQuAD dataset into Italian.
1 paper · 0 benchmarks
Synthetic Question Answering dataset in Serbian, acquired by automatic translation of SQuAD.
1 paper · 0 benchmarks
Confocal fluorescence microscopy is one of the most accessible and widely used imaging techniques for the study of biological processes at the cellular and subcellular levels.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
APPROVE consists of curated YouTube videos annotated with educational content.
1 paper · 1 benchmark
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
The Synthetic Signature Bankcheck Images (SSBI) Dataset is the first publicly available dataset of bank check images with annotations for detecting handwritten components, including names, amounts, dates, and signatures.
1 paper · 0 benchmarks
SSD_ID (Sub-Slot Dialogue dataset id number domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 0 benchmarks
SSD_NAME (Sub-Slot Dialogue dataset name domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 1 benchmark
SSD_PLATE (Sub-Slot Dialogue dataset license plate number domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is a dataset to benchmark real-time embedded object detection models for RoboCup SSL (Small Size League).
1 paper · 0 benchmarks
SSVC (Synthetic Structured Visual Content)
The Synthetic SVC (SSVC) dataset comprises 12,000 images with respective bounding box annotations and detailed graph representations.
1 paper · 0 benchmarks
We use Something-Something v2 dataset to obtain the generation prompts and ground truth masks from real action videos.
1 paper · 0 benchmarks
A large-scale Japanese video caption dataset consisting of 79,822 videos and 399,233 captions.
1 paper · 0 benchmarks
With the rise of social media, user-generated content has surged, and hate speech has proliferated.
1 paper · 0 benchmarks
STB (Stereo Hand Pose Benchmark)
3D hand pose data set created using stereo camera - contains 18,000 RGB images and paired depth images - 3D positions of hand joints (21 joints)
1 paper · 1 benchmark
STDW is a diverse large-scale dataset for table detection with more than seven thousand samples containing a wide variety of table structures collected from many diverse sources.
1 paper · 1 benchmark
StEduCov, a dataset annotated for stances toward online education during the COVID-19 pandemic.
1 paper · 1 benchmark
STEW (Simultaneous Task EEG Workload Dataset)
This dataset consists of raw EEG data from 48 subjects who participated in a multitasking workload experiment utilizing the SIMKAP multitasking test.
1 paper · 0 benchmarks
STN PLAD (STN Power Line Assets Dataset)
STN PLAD is a high-resolution and real-world image dataset of multiple high-voltage power line components.
1 paper · 1 benchmark
The STR-2021 dataset has 5,500 English sentence pairs manually annotated for semantic relatedness using a comparative annotation framework.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.