Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 119 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5665–5712 of 12,172
RemFX (RemFX Evaluation Datasets)
Audio samples processed with sound effects, to evaluate effect removal models.
3 papers · 0 benchmarks
Dataset Overview This dataset contains individual-level data from a randomized controlled trial (RCT) conducted in northern Uganda, along with associated satellite imagery.
3 papers · 0 benchmarks
ResQ (Real-world Spatial Question Answering)
ReSQ is a real-world Spatial Question Answering dataset with human-generated questions built on an existing corpus with SpRL annotations.
3 papers · 0 benchmarks
Researchy Questions is a set of about 100k Bing queries that users spent the most effort on.
3 papers · 0 benchmarks
Description 12-channel lung sounds for each patient Multi-channel Analysis opportunity 5 COPD severities (COPD0, COPD1, COPD2, COPD3, COPD4) Short-term recordings (At least 17s) The details: https://doi.org/10.28978/nesciences.349282 This…
3 papers · 1 benchmark
According to the WHO, World report on vision 2019, the number of visually impaired people worldwide is estimated to be 2.2 billion, of whom at least 1 billion have a vision impairment that could have been prevented or is yet to be…
3 papers · 1 benchmark
Retweet MTPP (Marked Temporal Point Processes on Retweet data)
This dataset contains time-stamped user retweet event sequences.
3 papers · 1 benchmark
Ricordi contains handwritten texts written in Italian.
3 papers · 0 benchmarks
Riedones3D is a dataset of 2,070 scans of coins.
3 papers · 0 benchmarks
RoFT is a dataset of 21,000 human annotations of generated text.
3 papers · 1 benchmark
RoboBEV is a robustness evaluation benchmark tailored for camera-based bird's eye view (BEV) perception under natural data corruptions and domain shift.
3 papers · 0 benchmarks
RoboPianist is a benchmarking suite for high-dimensional control, targeted at testing high spatial and temporal precision, coordination, and planning, all with an underactuated system frequently making-and-breaking contacts.
3 papers · 0 benchmarks
RobotPush is a dataset for object singulation – the task of separating cluttered objects through physical interaction.
3 papers · 0 benchmarks
Robotic Interestingness (Robotic Interestingness: A Dataset to Push the Limits of Online Visual Interesting Scene Prediction)
Robotic Interestingness dataset was created to promote the development visual interesting scene prediction for such purpose, for robots to better sense the world.
3 papers · 0 benchmarks
RuShiftEval is a manually annotated lexical semantic change dataset for Russian.
3 papers · 0 benchmarks
A dataset with high resolution (4K) images and manually-annotated dense labels every 50 frames.
3 papers · 0 benchmarks
A pool of real stocks from S&P 500 for recent 21 years from 01/02/2000 to 12/31/2020.
3 papers · 1 benchmark
S-VED (Sacrobosco Visual Element Dataset)
The Sacrobosco Visual Elements Dataset (S-VED) is derived from 359 Sphaera editions, centered on the Tractatus de sphaera by Johannes de Sacrobosco (—1256) and printed between 1472 and 1650.
3 papers · 0 benchmarks
SAVEE (Surrey Audio-Visual Expressed Emotion)
The Surrey Audio-Visual Expressed Emotion (SAVEE) dataset was recorded as a pre-requisite for the development of an automatic emotion recognition system.
3 papers · 1 benchmark
SCARED (Stereo Correspondence and Reconstruction of Endoscopic Data)
Sub-Challenge Part of the Endoscopic Vision Challenge
3 papers · 0 benchmarks
Student Classroom Behavior dataset (SCB-dataset) reflects real-life scenarios.
3 papers · 0 benchmarks
SCapRepo (Google Play Screenshot Caption)
A screenshot-caption dataset containing 135k pairs of screenshots and captions extracted from Google Play.
3 papers · 0 benchmarks
> High-quality underwater coral detection dataset for machine learning and computer vision research.
3 papers · 1 benchmark
We develop a primary dataset based on our task of suicide or depression classification.
3 papers · 0 benchmarks
SDD dataset contains a variety of indoor and outdoor scenes, designed for Image Defocus Deblurring.
3 papers · 1 benchmark
SDN (Situated Dialogue Navigation)
Situated Dialogue Navigation (SDN) is a navigation benchmark of 183 trials with a total of 8415 utterances, around 18.7 hours of control streams, and 2.9 hours of trimmed audio.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
Contains more than 6500 words semantically grouped under 110 categories.
3 papers · 0 benchmarks
A database of more than 2000 minutes of audio-visual data of 398 people coming from six cultures, 50% female, and uniformly spanning the age range of 18 to 65 years old.
3 papers · 0 benchmarks
scene graph labels for the 3D-FRONT dataset.
3 papers · 0 benchmarks
SHIFT15M is a dataset that can be used to properly evaluate models in situations where the distribution of data changes between training and testing.
3 papers · 0 benchmarks
A test dataset SICEGrad image datasets to represent complex mixed over-/under-exposed scenes.
3 papers · 1 benchmark
A test dataset SICEMix image datasets to represent complex mixed over-/under-exposed scenes.
3 papers · 1 benchmark
SIDD-Image (Segmented Intrusion Detection Dataset)
This is the first image-based network intrusion detection dataset.
3 papers · 1 benchmark
SILG (Symbolic Interactive Language Grounding)
Symbolic Interactive Language Grounding (SILG) is a multi-environment benchmark which unifies a collection of diverse grounded language learning environments under a common interface.
3 papers · 0 benchmarks
SKAB (Skoltech Anomaly Benchmark)
SKAB is designed for evaluating algorithms for anomaly detection.
3 papers · 2 benchmarks
SLNET (SLNET: A Redistributable Corpus of 3rd-party Simulink Models)
SLNET is collection of third party Simulink models.
3 papers · 0 benchmarks
SMILE-UHURA (Small Vessel Segmentation at Mesoscopic Scale from Ultra-High Resolution 7T Magnetic Resonance Angiogram)
The human brain receives nutrients and oxygen through an intricate network of blood vessels.
3 papers · 0 benchmarks
SMOKE (Real Dense Non-Uniform Fog)
The SMOKE dataset is a dataset for fog/smoke removal.
3 papers · 0 benchmarks
SONICS (Synthetic Or Not - Identifying Counterfeit Songs)
SONICS is a large-scale dataset comprising 97,164 songs — 48,090 real songs from YouTube and 49,074 fake songs from Suno & Udio — designed for synthetic song detection (SSD), also known as fake song detection (FSD).
3 papers · 0 benchmarks
SPEC5G is a dataset for the analysis of natural language specification of 5G Cellular network protocol specification.
3 papers · 0 benchmarks
A first-of-its-kind large dataset of sarcastic/non-sarcastic tweets with high-quality labels and extra features: (1) sarcasm perspective labels (2) new contextual features.
3 papers · 0 benchmarks
SPOT (Sentiment Polarity Annotations Dataset)
The SPOT dataset contains 197 reviews originating from the Yelp'13 and IMDB collections ([1][2]), annotated with segment-level polarity labels (positive/neutral/negative).
3 papers · 0 benchmarks
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models.
3 papers · 0 benchmarks
SYNTH-PEDES is a large-scale person dataset with image-text pairs by far, which contains 312,321 identities, 4,791,711 images, and 12,138,157 textual descriptions.
3 papers · 0 benchmarks
SaL-Lightning is a dataset for research in the field of Search as Learning.
3 papers · 0 benchmarks
SaRoCo is a dataset for detecting satire in Romanian news containing 55,608 news articles from multiple real and satirical news sources, of which 27,980 are regular and 27,628 satirical news reports.
3 papers · 0 benchmarks
The San Francisco Landmark Dataset contains a database of 1.7 million images of buildings in San Francisco with ground truth labels, geotags, and calibration data, as well as a difficult query set of 803 cell phone images taken with a…
3 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.