Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 209 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9985–10032 of 12,172
https://arxiv.org/abs/2503.15222
1 paper · 0 benchmarks
We present three defect rediscovery datasets mined from Bugzilla.
1 paper · 0 benchmarks
The full version of ReefSet used in Williams et al.
1 paper · 0 benchmarks
Ref-AVS seeks to segment objects within the visual domain based on expressions containing multimodal cues.
1 paper · 0 benchmarks
Refer360° is a novel large-scale referring expression recognition dataset consisting of 17,137 instruction sequences and ground-truth actions for completing these instructions in 360° scenes.
1 paper · 0 benchmarks
featureinput.npy Reflectance of 5 structures.
1 paper · 0 benchmarks
Teaching assistants (TAs) are heavily used in computer science courses as a way to handle high enrollment and still being able to offer students individual tutoring and detailed assessments.
1 paper · 0 benchmarks
RegDB-C is an evaluation set that consists of algorithmically generated corruptions applied to the RegDB test-set, and especially to both the visible and the thermal data.
1 paper · 0 benchmarks
This dataset was acquired in a retrospective study from a cohort of pediatric patients admitted with abdominal pain to Children’s Hospital St.
1 paper · 0 benchmarks
This is a dataset of regular expressions collected from regex101.com.
1 paper · 0 benchmarks
The primary environmental health threat in the WHO European Region is air pollution, impacting the daily health and well-being of its citizens significantly.
1 paper · 0 benchmarks
Dataset Details Total Labeled: 100% Labeled and Curated: 24,478 Pending: 0 Drafts: 0 Discarded: 696 High-Level Explanation This dataset includes labeled samples from the Colombian Aeronautical Regulations (RAC), covering all chapters…
1 paper · 0 benchmarks
This dataset is a collection of 5348 links from bug-introducing and bug-fixing commit sets extracted from Mozilla's Bugzilla with the use of bugbug.
1 paper · 0 benchmarks
The relational pattern similarity dataset is a new dataset upon the work of Zeichner et al.
1 paper · 0 benchmarks
The notebooks folder contains basic analysis of the organizational affiliation data for the contributors per open source project.
1 paper · 0 benchmarks
This dataset includes 3D point-cloud and 2D imagery from a flash LiDAR...
1 paper · 1 benchmark
RepLab 2013 dataset uses Twitter data in English and Spanish (more than 142,000 tweets).
1 paper · 0 benchmarks
Replay is a collection of multi-view, multi-modal videos of humans interacting socially.
1 paper · 0 benchmarks
Replication Data for: "DAM" (Replication Data for: "DAM: A Universal Dual Attention Mechanism for Multimodal Timeseries Cryptocurrency Trend Forecasting")
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Bitcoin is a peer-to-peer electronic payment system that popularized rapidly in recent years.
1 paper · 0 benchmarks
Blockchain has empowered computer systems to be more secure using a distributed network.
1 paper · 0 benchmarks
This dataset contains the data used for all statistical comparisons in our ICSV 2022 submission "Assessment of a Cost-Effective Headphone Calibration Procedure for Soundscape Evaluations", summarised in a single .csv file.
1 paper · 0 benchmarks
Decentralized finance (DeFi) is known for its unique mechanism design, which applies smart contracts to facilitate peer-to-peer transactions.
1 paper · 0 benchmarks
The dataset provides information about 450 HYIPs collected between November 2020 and September 2021.
1 paper · 0 benchmarks
Replication Data for: On estimating Armington elasticities for Japan's meat imports contains monthly import values and quantities from Jan 1996 to Dec 2020 for all 78 items.
1 paper · 0 benchmarks
As CryptoPunks pioneers the innovation of non-fungible tokens (NFTs) in AI and art, the valuation mechanics of NFTs has become a trending topic.
1 paper · 0 benchmarks
The model forecasts for the sub-seasonal forecasting application considered in the Online Learning under Optimism and Delay paper experiments.
1 paper · 0 benchmarks
This dataset contains the data used for all statistical analysis in our publication "Singapore Soundscape Site Selection Survey (S5): Identification of Characteristic Soundscapes of Singapore via Weighted k-means Clustering", summarised in…
1 paper · 0 benchmarks
Dataset and Stata codes for replicating Tables 1, 3 and Figures 1-4.
1 paper · 0 benchmarks
Census statistics play a key role in public policy decisions and social science research.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is our replication package for our study on Benchmarking scalability of stream processing frameworks deployed as microservices in the cloud.
1 paper · 0 benchmarks
This is the replication package for our systematic literature review and can be used for the reproducibility of the individual steps of our search and selection methodology.
1 paper · 0 benchmarks
This package contains the data and the reported results for the manuscript: Keo B, Li B, Younis W (2025) Measuring trade costs and analyzing the determinants of trade growth between Cambodia and major trading partners: 1993–2019.
1 paper · 0 benchmarks
RepoIMU T-stick The RepoIMU T-stick is a small, low-cost, and high-performance inertial measurement unit (IMU) that can be used for a wide range of applications.
1 paper · 0 benchmarks
This is a research artifact for the ICSE'22 paper "GitHub Sponsors: Exploring a New Way to Contribute to Open Source".
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LLM-generated output for compiling PDDL-domains and problems (5 scenarios, 5 runs per scenario): Specs 5 scenarios 5 trials.json - original json file Domain.pddl and Problem.pddl for each run are parsed out of the original json file.
1 paper · 0 benchmarks
A dataset to encourage the community to adapt oriented bounding box (OBB) detectors for more complex environments.
1 paper · 0 benchmarks
ReviewRobot Dataset Overview This repository contains data for paper ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis.
1 paper · 0 benchmarks
ata Set Name: Rice Dataset (Commeo and Osmancik) Abstract: A total of 3810 rice grain's images were taken for the two species (Cammeo and Osmancik), processed and feature inferences were made.
1 paper · 0 benchmarks
dataset (balanced) of 200 images consists of three classes - False Smut, Neck Blast and healthy grain class.
1 paper · 0 benchmarks
Riposte! (Riposte! A Large Corpus of Counter-Arguments)
From the Riposte!
1 paper · 0 benchmarks
Risholme-2021 contains >3.5K images of strawberries at various growth stages along with anomalous instances.
1 paper · 0 benchmarks
Risk-Aware Planning is a dataset that contains the overhead images and their semantic segmentation captured by a drone from the CityEnviron environment in AirSim simulator.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Ritter PoS (Ritter Twitter part-of-speech tagging)
PTB-tagged English Tweets
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.