Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 169 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8065–8112 of 12,172
Danish Airs and Grounds (DAG) is a large collection of street-level and aerial images targeting such cases.
1 paper · 0 benchmarks
A set of over 4000 dark corner artifact / vignette masks that can be applied to ISIC skin lesion images to test for the effect of such artifacts in classification tasks.
1 paper · 0 benchmarks
A pair Deblurring Benchmarking Dataset All scenes are collected in real-world scenarios.
1 paper · 0 benchmarks
This is a benchmark of data loss bugs for android apps.
1 paper · 0 benchmarks
IOPS and Latency measurements of a real data storage system
1 paper · 0 benchmarks
This archive contains raw data, intermediate results, statistics, and figures for the manuscript "Naïve individuals promote collective exploration in homing pigeons" Once unzipped, the folder structure will look as follow: - data/ [raw…
1 paper · 0 benchmarks
About the study This study was exploring the landscape of interpersonal conflicts during code review in following areas: - how these conflicts look like - what role do they play in software development - what are their consequences - what…
1 paper · 0 benchmarks
Data for "Image-based Backbone Reconstruction for Non-Slender Soft Robots" This dataset provides the data for the forthcoming paper "Image-based Backbone Reconstruction for Non-Slender Soft Robots".
1 paper · 0 benchmarks
Contains the current version of IMITATOR, all models and necessary scripts to reproduce all experiments on the benchmarks set.
1 paper · 0 benchmarks
Source: Linking Datasets on Organizations Using Half-a-Billion Open-Collaborated Records (Description (Markdown and LATEX enabled)) High-Level Explanation of the Dataset - Scale and Composition: This repository provides millions of…
1 paper · 0 benchmarks
For creating, optimizing, and evaluating our statistical model, we used the Public Unified Bug Dataset for Java.
1 paper · 0 benchmarks
The following experimental data were obtained on lithography devices made of magnetic multilayer tracks and thin tantalum transverse electrodes by Kerr microscopy and anomalous Hall effect measurements.
1 paper · 0 benchmarks
Data from the "Resistance Against Manipulative AI: key factors and possible actions" article
1 paper · 0 benchmarks
COVID-19 dataset for the world.
1 paper · 0 benchmarks
Data supporting: Improved Tangential Interpolation-based Multi-input Multi-output Modal Analysis of a Full Aircraft
1 paper · 0 benchmarks
DataCLUE is the first Data-Centric benchmark applied in NLP field.
1 paper · 0 benchmarks
Dataset Summary The DataSeeds.AI Sample Dataset (DSD) is a high-fidelity, human-curated computer vision-ready dataset comprised of 7,772 peer-ranked, fully annotated photographic images, 350,000+ words of descriptive text, and…
1 paper · 0 benchmarks
Testing results for the paper "Testing Theory of Mind in Large Language Models and Humans"
1 paper · 0 benchmarks
Simulation data and pre-trained Graph Neural Network (GNN) models produced in [1].
1 paper · 0 benchmarks
This is a part of dataset and models of the paper published in TMLR 2024 (Transactions on Machine Learning Research, https://jmlr.org/tmlr/).
1 paper · 0 benchmarks
Dataset information (e.g., google drive link) is attached in the GitHub repo: https://github.com/YY-GX/Annotated-Hands-Dataset Please find the description of the dataset in our paper: http://arxiv.org/abs/2401.15075
1 paper · 0 benchmarks
The dataset evaluates the number of vesicles observed in a Tcell that is close to a tumor cell.
1 paper · 0 benchmarks
Synthetic Datasets for ICSC Flagship 2.6.1.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is used for MPLP considering time windows constraints of customers and parking space.
1 paper · 0 benchmarks
This is a benchmark dataset for mid-price forecasting of limit order book data.
1 paper · 0 benchmarks
This dataset contains around 218K sentences, with 1.5 million words, from 30 different books designed for Post-OCR text correction.
1 paper · 0 benchmarks
Dataset for ZAugNet (Self-Supervised Z-Slice Augmentation for 3D Bio-Imaging via Knowledge Distillation)
Dataset used to train ZAugNet, a neural network for Z-slice augmentation, that encompasses a variety of shapes, textures, and microscopy techniques, as described below: Ascidian Embryos: This dataset consists of 3D confocal images of P.
1 paper · 0 benchmarks
Overview This dataset was collected during a pilot study that evaluated the virtual Cross Array Task (CAT) platform as an assessment tool for algorithmic thinking (AT) skills among K-12 students in Swiss compulsory education.
1 paper · 0 benchmarks
Data collected for the controlled experiment performed to analyze whether gamification can help in software testing education.
1 paper · 0 benchmarks
This dataset provides neutron and gamma-ray pulse signals for pulse shape discrimination experiments.
1 paper · 0 benchmarks
The dataset consists of three files: the metadata, comments, and captions of the ground-truth dataset videos collected and manually reviewed in this paper.
1 paper · 0 benchmarks
The Dataset contains more than 23500 3D garment models with their corresponding sewing patterns, each representing a unique garment design sampled from one of the 19 different categories.
1 paper · 0 benchmarks
We release both the processed data and evaluation results from our own experiments, and the underlying raw data that can be used for future experiments and schemes in the domain of Zero-Interaction Security.
1 paper · 0 benchmarks
This dataset of approximately 178,000 unique Dockerfiles collected from GitHub to facilitate sophisticated semantics-aware static analysis of Dockerfiles.
1 paper · 0 benchmarks
This Dataset contains the IDs of 5,427,024 commit authors who have created commits in git version control system, and have more than 1 ID in git.
1 paper · 0 benchmarks
Overview of the scoping review paper corpus, sorted by their diferent intent types, categories, and subcategories.
1 paper · 0 benchmarks
This data is for the Mis2-KDD 2021 under review paper: Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People’s Republic of China We present our dataset that focuses on propaganda techniques in Mandarin…
1 paper · 1 benchmark
This dataset contains 4,888 synthetic images of chess game states that occurred in games played by Magnus Carlsen.
1 paper · 0 benchmarks
Provides 450, 000 relevance annotations and 53 structured queries.
1 paper · 0 benchmarks
This is a part of dataset of the paper published in TMLR 2024 (Transactions on Machine Learning Research, https://jmlr.org/tmlr/).
1 paper · 0 benchmarks
This is a grounded dataset describing software-engineering problems in video-game development extracted from postmortems.
1 paper · 0 benchmarks
The dataset is generated from the study of computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
This repository contains the dataset for the study of the computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
This is the dataset to "Easing the Conscience with OPC UA: An Internet-Wide Study on Insecure Deployments" [In ACM Internet Measurement Conference (IMC ’20)].
1 paper · 0 benchmarks
Dataset outline This repository contains a novel time-series dataset for impact detection and localization on a plastic thin-plate, towards Structural Health Monitoring applications, using ceramic piezoelectric transducers (PZTs) connected…
1 paper · 0 benchmarks
Collected data from two distinct experiments in immersive, interactive VR where participants performed dynamic tasks as their eye, head, and hand movements were recorded.
1 paper · 0 benchmarks
This is the dataset used for classifying Gene-Disease relationship types from sentences.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.