Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 120 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5713–5760 of 12,172
Provides a set of stereo-rectified images and the associated groundtruthed disparities for 10 AOIs (Area of Interest) drawn from two sources: 8 AOIs from IARPA's MVS Challenge dataset and 2 AOIs from the CORE3D-Public dataset.
3 papers · 0 benchmarks
The ScandiQA dataset is a question-answering dataset specifically constructed for the Mainland Scandinavian languages, which include Danish, Norwegian, and Swedish.
3 papers · 0 benchmarks
Schiller contains handwritten texts written in modern German.
3 papers · 0 benchmarks
Schwerin contains handwritten texts written in medieval German.
3 papers · 0 benchmarks
SeaEval is a benchmark designed for evaluating multilingual foundation models (FMs).
3 papers · 0 benchmarks
SeaTurtleID is a public large-scale, long-span dataset with sea turtle photographs captured in the wild.
3 papers · 0 benchmarks
Dataset: The experiments are conducted using the Seaquest environment from the OpenAI Gym framework, which simulates the Atari 2600 game Seaquest.
3 papers · 1 benchmark
Aa new cross-season scaleless monocular depth prediction dataset from CMU Visual Localization dataset through structure from motion.
3 papers · 0 benchmarks
Secim2023 is a comprehensive dataset for social media researchers to study the upcoming election, develop tools to prevent online manipulation, and gather novel information to inform the public.
3 papers · 0 benchmarks
Given the unavailability of real-world pharmaceutical inspection-domain datasets, we have created the Sensum Solid Oral Dosage Forms (SensumSODF) dataset intended for research and evaluation purposes.
3 papers · 0 benchmarks
Data variables and description.
3 papers · 0 benchmarks
An annotated dataset of 161 episodes from three popular American TV serials: Breaking Bad, Game of Thrones and House of Cards.
3 papers · 0 benchmarks
Shmoop Corpus is a dataset of 231 stories that are paired with detailed multi-paragraph summaries for each individual chapter (7,234 chapters), where the summary is chronologically aligned with respect to the story chapter.
3 papers · 0 benchmarks
A short clip of video may contain progression of multiple events and an interesting story line.
3 papers · 3 benchmarks
SiW (Spoofing in the Wild) is a face anti-spoofing dataset recently introduced in [29] where images are extracted from short videos captured at high resolution and 30 frames per second.
3 papers · 1 benchmark
LA-2A Compressor data to accompany the paper "SignalTrain: Profiling Audio Compressors with Deep Neural Networks," https://arxiv.org/abs/1905.11928 Accompanying computer code: https://github.com/drscotthawley/signaltrain A collection of…
3 papers · 0 benchmarks
SketchHairSalon is a dataset for hair generation containing thousands of annotated hair sketch-image pairs and corresponding hair mattes.
3 papers · 0 benchmarks
SkinCon is a skin disease dataset densely annotated by dermatologists.
3 papers · 0 benchmarks
The SmartSpeaker benchmark tests the performance of reacting to music player commands in English as well as in French.
3 papers · 1 benchmark
The Specs on Faces (SoF) dataset, a collection of 42,592 (2,662×16) images for 112 persons (66 males and 46 females) who wear glasses under different illumination conditions.
3 papers · 0 benchmarks
The SoccerNet Game State Reconstruction task is a novel high level computer vision task that is specific to sports analytics.
3 papers · 0 benchmarks
Social Relation Dataset is a dataset for social relation trait prediction from face images.
3 papers · 0 benchmarks
Software Heritage is the largest existing public archive of software source code and accompanying development history.
3 papers · 0 benchmarks
Original SID dataset is introduced in "Learning to See in the Dark".
3 papers · 1 benchmark
Sound-Dr (Sound-Dr: Reliable Sound Dataset and Baseline Artificial Intelligence System for Respiratory Illnesses)
As the burden of respiratory diseases continues to fall on society worldwide, this paper proposes a high-quality and reliable dataset of human sounds for studying respiratory illnesses, including pneumonia and COVID-19.
3 papers · 0 benchmarks
These files are supplementary material for “Generalized Seismic Phase Detection with Deep Learning” by Ross et al.
3 papers · 1 benchmark
SpaGBOL (Spatial-Graph-Based Orientated Cross-View Geo-Localisation)
Cross-View Geo-Localisation within urban regions is challenging in part due to the lack of geo-spatial structuring within current datasets and techniques.
3 papers · 1 benchmark
An open source Multi-View Overhead Imagery dataset with 27 unique looks from a broad range of viewing angles (-32.5 degrees to 54.0 degrees).
3 papers · 0 benchmarks
Spanish TimeBank 1.0 was developed by researchers at Barcelona Media and consists of Spanish texts in the AnCora corpus annotated with temporal and event information according to the TimeML specification language.
3 papers · 1 benchmark
Description A dataset of assembly functions that are vulnerable to Spectre-V1 attack.
3 papers · 0 benchmarks
The dataset is approved for public release, distribution unlimited.
3 papers · 0 benchmarks
Spoken versions of the Semantic Textual Similarity dataset for testing semantic sentence level embeddings.
3 papers · 0 benchmarks
Stanceosaurus is a corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.
3 papers · 0 benchmarks
In recent years, the number of range scanners and surface reconstruction algorithms has been growing rapidly.
3 papers · 0 benchmarks
This dataset contains a set of sentences by extracting all the sentences mentioning the term from the court decisions retrieved from the Caselaw access project data.
3 papers · 0 benchmarks
This dataset comprises a collection of stellarator configurations used to train the model over multiple iterations.
3 papers · 0 benchmarks
This repository contains a financial-domain-focused dataset for financial sentiment/emotion classification and stock market time series prediction.
3 papers · 0 benchmarks
A new dataset for streaming classification consisting of temporally correlated images from 51 distinct object categories and additional evaluation classes outside of the training distribution to test novelty recognition.
3 papers · 0 benchmarks
Presents a new dataset of code snippets with short descriptions, created using data gathered from Stackoverflow, a popular programming help website.
3 papers · 0 benchmarks
Introduction This dataset supports Ye et al.
3 papers · 0 benchmarks
SupplyGraph (SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks)
Graph Neural Networks (GNNs) have gained traction across different domains such as transportation, bio-informatics, language processing, and computer vision.
3 papers · 0 benchmarks
SurgT is a dataset for benchmarking 2D Trackers in Minimally Invasive Surgery (MIS).
3 papers · 0 benchmarks
A personalized subset of Symbolic Mathematics dataset, initially introduced in the paper Deep Learning for Symbolic Mathematics (Lample et al.).
3 papers · 0 benchmarks
Synthetic dataset for polyp segmentation.
3 papers · 0 benchmarks
Synthehicle is a massive CARLA-based synthehic multi-vehicle multi-camera tracking dataset and includes ground truth for 2D detection and tracking, 3D detection and tracking, depth estimation, and semantic, instance and panoptic…
3 papers · 1 benchmark
TCAB (Text Classification Attack Benchmark)
Text Classification Attack Benchmark (TCAB) is a dataset for analyzing, understanding, detecting, and labeling adversarial attacks against text classifiers.
3 papers · 0 benchmarks
Contains static tasks as well as a multitude of more dynamic tasks, involving larger motion of the hands.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.