Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 142 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6769–6816 of 12,172
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
RHM (Rhm: Robot house multi-view human activity recognition dataset)
The Robot House Multi-View dataset (RHM) contains four views: Front, Back, Ceiling, and Robot Views.
2 papers · 1 benchmark
RIR dataset (Planar Room Impulse Response Dataset - ACT, DTU Electro (b. 355 r. 008))
Dataset of Room Impulse Responses measured at the Acoustic Technology group facilities, DTU Electro.
2 papers · 0 benchmarks
RISE is a large-scale video dataset for Recognizing Industrial Smoke Emissions.
2 papers · 0 benchmarks
RISEdb (Robust Indoor Localization in Complex Scenarios (RISE) database)
The RISE (Robust Indoor Localization in Complex Scenarios) dataset is meant to train and evaluate visual indoor place recognizers.
2 papers · 0 benchmarks
RLAP (Remote Learning Affect and Physiologic dataset)
The Remote Learning Affect and Physiologic (RLAP) dataset is a dataset applied to remote learning affect and engagement, which contains learners' blood volume pulse (BVP) signals that are highly synchronized.
2 papers · 0 benchmarks
RLD (Responsive Listener Dataset)
RLD (Responsive Listener Dataset) is a conversation video corpus collected from the public resources featuring 67 speakers, 76 listeners with three different attitudes.
2 papers · 0 benchmarks
RL Unplugged is suite of benchmarks for offline reinforcement learning.
2 papers · 0 benchmarks
The first benchmark comprising 473 prompts designed to assess the ability of LLMs to resist malicious code generation.
2 papers · 0 benchmarks
An environment for RNA design given structure constraints with structures from different datasets to choose from.
2 papers · 0 benchmarks
ROPE (Recognition-based Object Probing Evaluation)
We introduce Recognition-based Object Probing Evaluation (ROPE), an automated evaluation protocol that considers the distribution of object classes within a single image during testing and uses visual referring prompts to eliminate…
2 papers · 0 benchmarks
RSDD-Time is a dataset of 598 manually annotated self-reported depression diagnosis posts from Reddit that include temporal information about the diagnosis.
2 papers · 0 benchmarks
RSOC (Remote Sensing Object Counting)
RSOC is a large-scale object counting dataset with remote sensing images, which contains four important geographic objects: buildings, crowded ships in harbors, large-vehicles and small-vehicles in parking lots.
2 papers · 0 benchmarks
RTC is a benchmark corpus of social media comments sampled over three years.
2 papers · 0 benchmarks
RUFF is a large-scale dataset to measure pronoun fidelity in English.
2 papers · 0 benchmarks
RUGD (RUGD: Robot Unstructured Ground Driving)
A Video Dataset for Visual Perception and Autonomous Navigation in Unstructured Environments.
2 papers · 1 benchmark
RUSLAN is a Russian spoken language corpus for text-to-speech task.
2 papers · 0 benchmarks
The RaDelft dataset is a novel, large-scale, real-life, and multi-sensor dataset that has been recorded using a demonstrator vehicle in different locations in the city of Delft.
2 papers · 0 benchmarks
Abstract In the view of national security, radar micro-Doppler (m-D) signatures-based recognition of suspicious human activities becomes significant.
2 papers · 2 benchmarks
RaidaR (RaidaR: A Rich Annotated Image Dataset of Rainy Street Scenes)
RaidaR is a rich annotated image dataset of rainy street scenes.
2 papers · 0 benchmarks
RainNet is a real (non-simuated) large-scale spatial precipitation downscaling dataset that contains 62,424 pairs of low-resolution and high-resolution precipitation maps for 17 years.
2 papers · 0 benchmarks
RaindropClarity (A Dual-Focused Dataset for Day and Night Raindrop Removal)
Existing raindrop removal datasets have two shortcomings.
2 papers · 0 benchmarks
A dataset consisting of recipient 46 users and, 26180 tweets.
2 papers · 0 benchmarks
Dataset used in the publication of Rapid Design of Top-Performing Metal-Organic Frameworks with Qualitative Representations of Building Blocks.
2 papers · 0 benchmarks
Data annotation The 1,073 full rare disease mention annotations (from 312 MIMIC-III discharge summaries) are in fullsetRDannMIMICIIIdisch.csv.
2 papers · 1 benchmark
This dataset consists of an unpaired and paired set of images captured by two different smartphone cameras: Samsung Galaxy S9 and iPhone X.
2 papers · 0 benchmarks
Our dataset consists of over 1000 fractured frescoes.
2 papers · 0 benchmarks
RePack is a dataset to study the detection of repackaged Android apps.
2 papers · 0 benchmarks
ReactionGIF is an affective dataset of 30K tweets which can be used for tasks like induced sentiment prediction and multilabel classification of induced emotions.
2 papers · 0 benchmarks
Real-CE is a real-world Chinese-English benchmark dataset for the task of STISR with the emphasis on restoring structurally complex Chinese characters.
2 papers · 0 benchmarks
The dataset contains patches of facial reflectance as described in the paper, namely the diffuse albedo, diffuse normals, specular albedo, specular normals, as well as the shape in UV space.
2 papers · 0 benchmarks
RedEval is a safety evaluation benchmark designed to assess the robustness of large language models (LLMs) against harmful prompts.
2 papers · 0 benchmarks
The C-SSRS dataset contains 500 Reddit posts from the subreddit r/depression.
2 papers · 0 benchmarks
Articles originating from subreddits with explicitly stated ideologies are categorized into three groups: 72,488 articles in the Liberal class, 79,573 articles in the Conservative class, and 225,083 articles in the Restricted class.
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Rendered Handpose Dataset contains 41258 training and 2728 testing samples.
2 papers · 0 benchmarks
The Rendered SST2 dataset is a dataset released by OpenAI, that measures the optical character recognition capability of visual representations.
2 papers · 1 benchmark
Rent3D++ is an extension of the Rent3D floorplans + photos dataset.
2 papers · 1 benchmark
ReplicaGrasp dataset is created by spawning objects from GRAB into the ReplicaCAD scenes, simulated in random positions and orientations using the Habitat simulator.
2 papers · 0 benchmarks
Transaction fee mechanism (TFM) is an essential component of a blockchain protocol.
2 papers · 0 benchmarks
The dataset consists of three files: a file with behaviour data (events.csv), a file with item properties (itemproperties.сsv) and a file, which describes category tree (categorytree.сsv).
2 papers · 1 benchmark
The Retina Benchmark is a set of real-world tasks that accurately reflect such complexities and are designed to assess the reliability of predictive models in safety-critical scenarios.
2 papers · 0 benchmarks
The Retinal Microsurgery dataset is a dataset for surgical instrument tracking.
2 papers · 0 benchmarks
Noiseless reverberant dataset using the public WSJ0 corpus and simulated room impulse responses using the PyRoomAcoustics library.
2 papers · 0 benchmarks
ReviewQA is a question-answering dataset based on hotel reviews.
2 papers · 0 benchmarks
The Rhythmic Gymnastics dataset contains videos of four different types of gymnastics routines: ball, clubs, hoop and ribbon.
2 papers · 1 benchmark
We collect a dataset of Rich Human Feedback on 18K images (RichHF-18K), which contains (i) point annotations on the image that highlight regions of implausibility/artifacts, and text-image misalignment; (ii) labeled words on the prompts…
2 papers · 0 benchmarks
RoFT-chatgpt is a variation of RoFT dataset, where the same human prompts are continued with the gpt-3.5-turbo model.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.