Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 208 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9937–9984 of 12,172
Rogue Wave Dataset-10K dataset consists of 10191 rogue wave images.
1 paper · 0 benchmarks
We create Rwanda built-up regions dataset, a different and versatile in nature from previously available datasets.
1 paper · 0 benchmarks
The data set contains multimodal sensor data generated by a tracked mobile robot in an outdoor and an indoor environemnt.
1 paper · 0 benchmarks
To effectively measure the alignment between automatic evaluation metrics and radiologists' assessments in medical text generation tasks, we have established a comprehensive benchmark, RaTE-Eval, that encompasses three tasks, each with its…
1 paper · 0 benchmarks
RaTE-NER dataset is a large-scale, radiological named entity recognition (NER) dataset, including 13,235 manually annotated sentences from 1,816 reports within the MIMIC-IV database, that spans 9 imaging modalities and 23 anatomical…
1 paper · 0 benchmarks
Annotated Earth Observation dataset of extreme events
1 paper · 0 benchmarks
(WBC) dataset which consisted of 14514 WBC images across five classes 301 basophils, 795 monocytes, 1066 eosinophils, 8891 neutrophils, and 3461 lymphocytes at resolutions of 575 x 575.
1 paper · 0 benchmarks
http://hf.co/datasets/govtech/rabakbench
1 paper · 0 benchmarks
A fully-annotated, open-design dataset of autonomous and piloted high-speed flight
1 paper · 0 benchmarks
RadCases Dataset This HuggingFace (HF) dataset contains the raw case labels for input patient "one-liner" case summaries according to the ACR Appropriateness Criteria.
1 paper · 0 benchmarks
The ROAD dataset is made up of observations from the Low Frequency Array (LOFAR) telescope.
1 paper · 0 benchmarks
A total of 227 cross sectional images (20 x 54 mm with a resolution of 289 x 648 pixels) of hind-leg xenograft tumors from 29 mice were obtained with 1mm step-wise movement of the array mounted on a manual positioning device.
1 paper · 0 benchmarks
A multimodal dataset of radio galaxies and their corresponding infrared hosts.
1 paper · 0 benchmarks
Automating the creation of catalogues for radio galaxies in next-generation deep surveys necessitates the identification of components within extended sources and their respective infrared hosts.
1 paper · 1 benchmark
RadioTalk is a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019.
1 paper · 0 benchmarks
RaindropsOnWindshield is a dataset for training and assessing vision algorithms' performance for different tasks of image artifacts detection on either camera lens or windshield.
1 paper · 0 benchmarks
Versatile synthetic classification dataset based on precise input spike timings drawn from smooth random manifolds as previously described
1 paper · 0 benchmarks
The dataset contains generated random signals for autoencoding purposes.
1 paper · 0 benchmarks
This is a repository of PCAP files obtained by executing ransomware binaries and capturing the network traffic created when encrypting a set of files shared from an SMB server.
1 paper · 0 benchmarks
The NMR-POISE paper can be found at: Anal.
1 paper · 0 benchmarks
This dataset contains raw data of baseline experimental results.
1 paper · 0 benchmarks
Raw-Microscopy: 940 raw bright-field microscopy images of human blood smear slides for leukocyte classification (microscopy/images/rawscale100) with corresponding labels (microscopy/labels).
1 paper · 0 benchmarks
RawNIND (Raw Natural Image Noise Dataset)
The Raw Natural Image Noise Dataset (RawNIND) is a diverse collection of paired raw images designed to support the development of denoising models that generalize across sensors, image development workflows, and styles.
1 paper · 0 benchmarks
RawRipe Dataset (Fruit Maturity Recognition from Agricultural, Market and Automation Perspectives, IECON'21)
The dataset comprises of images of 10 types of fruits in both raw and ripe states.
1 paper · 2 benchmarks
Data obtained by ray-tracing simulation in Herald Square.
1 paper · 0 benchmarks
realfred is an embodied instruction following benchmark.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Efficiently rescaling a large dataset by adapting statistical computation to validation results of a pre-trained network.
1 paper · 0 benchmarks
Two-person interaction dataset consisting of fullbody and hand motions.
1 paper · 0 benchmarks
Synthetic humans generated by the RePoGen method.
1 paper · 0 benchmarks
Open KB canonicalization dataset.
1 paper · 0 benchmarks
Two versions of the dataset are offered: one is the full dataset used to train the models in our paper, and the other is a mini dataset for easier examination.
1 paper · 0 benchmarks
A genomics dataset for OOD detection that allows other researchers to benchmark progress on this important problem.
1 paper · 0 benchmarks
This dataset represents residential real estate listings with the following features: ZIP: The ZIP code where the property is located || SOLDPRICE: The listing price of the property in USD || SQFT: The square footage (living area) of the…
1 paper · 0 benchmarks
A total of 80 real material samples were captured in a dark room.
1 paper · 0 benchmarks
Presented data contains the record of five spreading campaigns that occurred in a virtual world platform.
1 paper · 0 benchmarks
A real-world stereo video dataset, containing 1200 frame pairs with real-world color and sharpness mismatches caused by beam splitter.
1 paper · 0 benchmarks
The ground truth betweenness-centralities for the real-world graphs are provided by AlGhamdi et al.
1 paper · 0 benchmarks
This is the official dataset collected for to test the sim-to-real transfer.
1 paper · 0 benchmarks
This repository contains the realFormula dataset presented in the paper MathNet: A Data-Centric Approach for Printed Mathematical Expression Recognition.
1 paper · 0 benchmarks
RealHDRTV dataset is the first real-world paired SDRTV-HDRTV dataset, which includes SDRTV-HDRTV pairs with 8K resolutions captured by a smartphone camera with the “SDR” and “HDR10” modes.
1 paper · 0 benchmarks
RealVul (RealVul-Vulnerability Dataset following realistic settings)
This is a C++ vulnerability detection dataset following realistic settings.
1 paper · 0 benchmarks
The Reasonable Crowd dataset is a dataset to evaluate autonomous driving in a limited operating domain.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Reddit Engagement Dataset (RED), a distant-supervision set, with 80k single-turn conversations.
1 paper · 0 benchmarks
Dataset with articles posted in the r/Liberal and r/Conservative subreddits.
1 paper · 1 benchmark
This is a dataset of over 40K Reddit comments removed by moderators according to the specific type of macro norm being violated.
1 paper · 0 benchmarks
This dataset comprises 77,175 Reddit posts from 115 subreddit forums, annotated for the presence of 15 topics related to eating disorders and dieting.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.