Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 207 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9889–9936 of 12,172
In the actual globalized world, the transportation of goods between any country is something normal.
1 paper · 0 benchmarks
RASMD (RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions)
Current autonomous driving algorithms heavily rely on the visible spectrum, which is prone to performance degradation in adverse conditions like fog, rain, snow, glare, and high contrast.
1 paper · 0 benchmarks
RB-Dust (RB-Dust: Real-world Industrial Dust Dehazing Dataset)
A small-scale real-world dataset containing hazy/dusty industrial images and their clean ground truth counterparts.
1 paper · 1 benchmark
We conducted a large crowdsourcing study of click patterns in an interactive segmentation scenario and collected 475K real-user clicks.
1 paper · 0 benchmarks
REAP is a digital benchmark that allows the user to evaluate patch attacks on real images, and under real-world conditions.
1 paper · 0 benchmarks
Scene Text Recognition training data
1 paper · 0 benchmarks
REBUS (A Robust Evaluation Benchmark of Understanding Symbols)
Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input.
1 paper · 1 benchmark
Relations in Captions (REC-COCO) is a new dataset that contains associations between caption tokens and bounding boxes in images.
1 paper · 0 benchmarks
REFCAT (Internet Archive Scholar Reference Dataset)
Internet Archive Scholar Reference Dataset.
1 paper · 0 benchmarks
Reader Emotion News 20k Dataset
1 paper · 0 benchmarks
RES-Q (RES-Q: Evaluating Code-Editing Large Language Model Systems at the Repository Scale)
RES-Q is a natural language instruction-based benchmark for evaluating Repository Editing Systems, which consists of 100 handcrafted repository editing tasks derived from real GitHub commits.
1 paper · 1 benchmark
Asthma is a common, usually long-term respiratory disease with negative impact on society and the economy worldwide.
1 paper · 0 benchmarks
This dataset is used for RF signal recognition, used to recognize different RF devices based on the signals they transmitted.
1 paper · 0 benchmarks
RFSD (Russian Financial Statements Database)
The Russian Financial Statements Database (RFSD) The Russian Financial Statements Database (RFSD) is an open, harmonized collection of annual unconsolidated financial statements of the universe of Russian firms.
1 paper · 0 benchmarks
In this paper, we propose RFUAV as a new benchmark dataset for radio-frequency based (RF-based) unmanned aerial vehicle (UAV) identification and address the following challenges: Firstly, many existing datasets feature a restricted variety…
1 paper · 0 benchmarks
RGB Arabic Alphabet Sign Language (AASL) dataset
1 paper · 1 benchmark
RGRS (ResearchGate dataset for Recommending Systems)
RGRS is a dataset for collaboratior recommendation on the ResearchGate academic social network.
1 paper · 0 benchmarks
The data used in - "Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy" (Bowles et al.
1 paper · 0 benchmarks
The RIKEN Microstructural Imaging Metadatabase is a semantic web-based imaging database in which image metadata are described using the Resource Description Framework (RDF) and detailed biological properties observed in the images can be…
1 paper · 0 benchmarks
This database consists of two main components; data on COVID-19 infections and data on the ICU occupancy of COVID-19 patients.
1 paper · 0 benchmarks
The datasets of "Reinforcement Learning-enhanced Shared-account Cross-domain Sequential Recommendation" (TKDE 2022)
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A benchmark suite of continuous control tasks, including classic tasks like cart-pole swing-up, tasks with very high state and action dimensionality such as 3D humanoid locomotion, tasks with partial observations, and tasks with…
1 paper · 0 benchmarks
RLM25 (Research-Level Mathematics 2025)
RLM25 is an evaluation benchmark containing 619 paired examples of research-level natural language mathematical statements and their corresponding Lean formalizations.
1 paper · 0 benchmarks
In this dataset, various objects are arranged on a white table.
1 paper · 0 benchmarks
The RMRC 2014 indoor dataset is a dataset for indoor semantic segmentation.
1 paper · 0 benchmarks
ROAST (Review level Opinion Aspect Sentiment Target Joint Detection for ABSA)
This repository has a review-level multidomain multilingual dataset for Aspect-based Sentiment Analysis(ABSA) for the paper ROAST: Review-level Opinion Aspect Sentiment Target Joint Detection.
1 paper · 0 benchmarks
Radiology Objects in COntext (ROCO): A Multimodal Image Dataset
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
RPCD (Reddit Photo Critique Dataset)
The Reddit Photo Critique Dataset (RPCD) contains tuples of image and photo critiques.
1 paper · 0 benchmarks
RPEval (Role-Playing Evaluation Dataset)
Role-Playing Eval (RPEval) is a benchmark dataset designed to evaluate large language models' role-playing abilities across emotional understanding, decision-making, moral alignment, and in-character consistency.
1 paper · 0 benchmarks
RRG (Russian RST dataset from GUM v9.1 corpus)
Parallel version of annotations in GUM RST v9.1.
1 paper · 0 benchmarks
The following files contains the simulation inputs and outputs for conducting the multi-objetive optimization of thermal comfort and dyalight with the Response Surface Methodology.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Pre-rendered dataset used in Training and Predicting Visual Error for Real-Time Applications for the Emerald Square scenes.
1 paper · 0 benchmarks
Pre-rendered dataset used in Training and Predicting Visual Error for Real-Time Applications for the Lumberyard Bistro scenes.
1 paper · 0 benchmarks
Pre-rendered dataset used in Training and Predicting Visual Error for Real-Time Applications for the Sibenik Cathedral scene.
1 paper · 0 benchmarks
Pre-rendered dataset used in Training and Predicting Visual Error for Real-Time Applications for the Sun Temple scene.
1 paper · 0 benchmarks
RTB (Robot Tracking Benchmark)
The Robot Tracking Benchmark (RTB) is a synthetic dataset that facilitates the quantitative evaluation of 3D tracking algorithms for multi-body objects.
1 paper · 1 benchmark
RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).
1 paper · 0 benchmarks
A corpus of real-world spoken personal narratives comprising 10,296 narrative clauses from 594 video transcripts.
1 paper · 0 benchmarks
Multilingual explainable fact-checking dataset on Russia-Ukraine Conflict 2022
1 paper · 0 benchmarks
RUSHOLD (Roman Urdu Hate Speech and Offensive Language Dataset)
RUHSOLD is hate speech and offensive language dataset in Roman Urdu.
1 paper · 0 benchmarks
RUSS (Rapid Universal Support Service) is a dataset that consists of a collection of 741 real-world step-by-step natural language instructions (raw and annotated) from the open web, and for each: its corresponding webpage DOM, ground-truth…
1 paper · 0 benchmarks
RVL-CDIPMP is our first contribution to retrieve the original documents of the IIT-CDIP test collection which were used to create RVL-CDIP.
1 paper · 0 benchmarks
RVL-CDIPMP-N can serve its original goal as a covariate shift test set, now for multi-page document classification.
1 paper · 0 benchmarks
The RWCP Sound Scene Database includes non-speech sounds recorded in an anechoic room, reconstructed signals in various rooms, impulse responses for a microphone array, speech data recorded with the same array, and recordings of background…
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.