Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 210 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10033–10080 of 12,172
Dataset contains light curves of 6 rocket body types from Mini Mega Tortora database (MMT)[^1].
1 paper · 0 benchmarks
A large dataset from games of some of the top teams (from 2016 and 2017) in RoboCup Soccer Simulation League (2D), where teams of 11 robots (agents) compete against each other.
1 paper · 0 benchmarks
Robot@Home2 (Robot@Home2, a robotic dataset of home environments)
Robot@Home2, is an enhanced version aimed at improving usability and functionality for developing and testing mobile robotics and computer vision algorithms.
1 paper · 0 benchmarks
Robust Summarization Evaluation Benchmark is a large human evaluation dataset consisting of over 22k summary-level annotations over state-of-the-art systems on three datasets.
1 paper · 0 benchmarks
This synthetic event dataset is used in Robust e-NeRF to study the collective effect of camera speed profile, contrast threshold variation and refractory period on the quality of NeRF reconstruction from a moving event camera.
1 paper · 0 benchmarks
Tagged for Sentiment (Positive, Negative, Neutral).
1 paper · 0 benchmarks
The Room environment - v0 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
The Room environment - v1 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
The Room environment - v2 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
RoomSpace: a new benchmark designed to evaluate language models on spatial reasoning tasks demanding spatial relation knowledge and multi-hop reasoning.
1 paper · 0 benchmarks
The RoseBlooming dataset is a stage-specific flower dataset for detection.
1 paper · 0 benchmarks
Fact-based Text Editing dataset based on RotoWire dataset
1 paper · 1 benchmark
The RotoWire-Modified dataset is a cleaned extension of the RotoWire dataset, with writer information about each document.
1 paper · 0 benchmarks
RuDaS (Synthetic Datasets for Rule Learning)
Logical rules are a popular knowledge representation language in many domains.
1 paper · 1 benchmark
https://github.com/dialogue-evaluation/RuOpinionNE-2024
1 paper · 0 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
The work provides a comprehensive overview of the corpus for the Russian language for the commonsense inference task.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
S-BIAD843 (Individual 3D cell shapes of Drosophila Wing Disc)
Late third instar wing imaginal discs were cultured in Shields and Sang M3 media (Sigma) supplemented with 2% FBS (Sigma), 1% pen/strep (Gibco), 3ng/ml ecdysone (Sigma) and 2ng/ml insulin (Sigma).
1 paper · 0 benchmarks
S-ODv2 (SeaDronesSee-Object Detection v2)
SeaDronesSee-Object Detection v2 (S-ODv2) dataset contains 14,227 RGB images (training: 8,930; validation: 1,547; testing: 3,750).
1 paper · 0 benchmarks
S-SOD (Surveillance Salient Object Detection)
To validate the generalization abilities of SOD models, we create a small-scale dataset by collecting the most challenging images with varying brightness and contrast, background and foreground colors overlap, among many others.
1 paper · 0 benchmarks
S2B (Symbolic Behaviour Benchmark)
Suite of OpenAI Gym-compatible multi-agent reinforcement learning environment centered around meta-referential games to benchmark for behavioral traits pertaining to symbolic behaviours, as described in Santoro et al., 2021, "Symbolic…
1 paper · 0 benchmarks
S3O4D (Stanford 3D Objects for Disentangling)
The data consists of 100,000 renderings each of the Bunny and Dragon objects from the Stanford 3D Scanning Repository.
1 paper · 0 benchmarks
SACID (Saliency Aware Compressed Images Dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SAD-Instruct (Situational Awareness Database for Instruct-Tuning)
The Situational Awareness Database for Instruct-Tuning (SAD-Instruct) is a dataset for dynamic task guidance.
1 paper · 0 benchmarks
SAGC-A68 (A space access graph dataset for the classification of spaces and space elements in apartment buildings)
The analysis of building models for usable area, building safety, and energy efficiency requires accurate classification data of spaces and space elements.
1 paper · 0 benchmarks
SAIL 2017 (Sentiment Analysis for Indian Languages)
India is a linguistic area with one of the longest histories of contact, influence, use, teaching and learning of English-in-diaspora in the world (Kachru and Nelson, 2006).
1 paper · 1 benchmark
SALAMI (Structural Analysis of Large Amounts of Music Information)
Comes from https://ddmal.music.mcgill.ca/research/SALAMI/: SALAMI is an innovative and ambitious computational musicology project.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
Sara motion is a 3D motion dataset, named Synthetic Actors and Real Actions (SARA), for training a model to produce motion embeddings suitable for reasoning about motion similarity.
1 paper · 0 benchmarks
Doi: 10.1101/2020.04.24.20078584
1 paper · 0 benchmarks
SAS-Bench represents the first specialized benchmark for evaluating Large Language Models (LLMs) on Short Answer Scoring (SAS) tasks.
1 paper · 0 benchmarks
SAT-MTB-VSR is a large-scale dataset for satellite video super-resolution made from original videos of Jilin-1, which is a subset of the satellite video multitasking dataset SAT-MTB.
1 paper · 1 benchmark
SBA (Sequentail Brick Assembly Dataset)
The RAD (Randomly Assembled Object Construction) dataset is a synthetic 3D LEGO dataset designed for the task of Sequential Brick Assembly (SBA).
1 paper · 0 benchmarks
The SBCoseg dataset includes 889 groups of images and each group consists of 18 images with a common object, leading to 16002 images in total.
1 paper · 1 benchmark
Building a large-scale figure QA dataset requires a considerable amount of work, from gathering and selecting figures to extracting attributes like text, numbers, and colors, and generating QAs.
1 paper · 0 benchmarks
SBU-WSD-Corpus is a corpus for Persian Word Sense Disambiguation (WSD).
1 paper · 0 benchmarks
SC2EGSet: StarCraft II Esport Game State Dataset Pre-processed data that was generated from the SC2ReSet: StarCraft II Esports Replaypack Set Data Modeling Our aplication programing interface (API) implementation supports downloading,…
1 paper · 0 benchmarks
Raw StarCraft II data is subject to processing under the Blizzard end user license agreement (EULA), and in special cases Blizzard AI and Machine Learning License may be applied.
1 paper · 0 benchmarks
The dataset SCARED-C is introduced in the context of assessing robustness in endoscopic depth prediction models.
1 paper · 1 benchmark
SCG (SCG Dataset from Graph Neural Networks in Supply Chain Analytics and Optimization: Concepts, Perspectives, Dataset & Benchmarks)
Abstract: Graph Neural Networks (GNNs) have recently gained traction in transportation, bioinformatics, language and image processing, but research on their application to supply chain management remains limited.
1 paper · 1 benchmark
SCI (Self-Contradictory Instructions)
Large multimodal models (LMMs) excel in adhering to human instructions.
1 paper · 0 benchmarks
SCIAN (SCIAN Gold-standard for Morphological Sperm Analysis)
Dataset of sperm head images with expert-classification labels.
1 paper · 0 benchmarks
SCIMAT is a large question-answer dataset for mathematics and science problems; such dataset can have impact on online education, intelligent tutoring and automated grading.
1 paper · 0 benchmarks
SCMD2016 (Satellite Cloudage Map Dataset)
SCMD dataset is a brand new cloudage nowcasting dataset for deep learning research.
1 paper · 0 benchmarks
A Chinese sign language dataset that includes dialogue information.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.