Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 151 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7201–7248 of 12,172
This is a dataset for anomalies detection in 3D printing.
1 paper · 0 benchmarks
3D-Speaker is a large-scale speech corpus designed to facilitate the research of speech representation disentanglement.
1 paper · 0 benchmarks
3D-ZeF (3D ZebraFish Tracking Benchmark)
3D-ZeF dataset consists of eight sequences with a duration between 15-120 seconds and 1-10 free moving zebrafish.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
3DO Dataset | On the Generalization of WiFi-based Person-centric Sensing in Through-Wall Scenarios The 3DO dataset comprises 42 five-minute recordings (~1.25M WiFi packets) of three human activities performed by a single person, captured…
1 paper · 0 benchmarks
TDW is a 3D virtual world simulation platform, utilizing state-of-the-art video game engine technology.
1 paper · 0 benchmarks
3U-VQA (Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset)
To tackle the challenge of obtaining out-of-distribution (OOD) data for LVQA models, we introduced a novel dataset named 3U-VQA dataset (Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset).
1 paper · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this DIF file.
1 paper · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this Excel file.
1 paper · 0 benchmarks
The 42Street dataset is based on a theater play as an example of such an application.
1 paper · 0 benchmarks
It contains 900 audio clips, annotated into 4 quadrants, according to Russell's model.
1 paper · 0 benchmarks
Every decade following the Census, states and municipalities must redraw districts for Congress, state houses, city councils, and more.
1 paper · 0 benchmarks
These are larger MATLAB .mat files required for reproducing plots from the sgbaird-5DOF/interp repository for grain boundary property interpolation.
1 paper · 0 benchmarks
This dataset contains two types of intercepted network packets: "normal" network traffic packets (i.e.
1 paper · 1 benchmark
The dataset contains 60,000 Stack Overflow questions from 2016-2020, classified into three categories: 1.
1 paper · 1 benchmark
6IMPOSE (Synthetic RGBD dataset for 6D pose estimation)
The dataset includes the synthetic data generated from rendering the 3D meshes of LM objects and several household objects in Blender for training 6D pose estimation algorithms.
1 paper · 0 benchmarks
This paper constructs 7-digit product Supply-Use Tables (SUTs) and symmetric Input-Output Tables (IOTs) for the Indian economy using microdata from the Annual Survey of Industries (ASI) for the period 2016-2021.
1 paper · 0 benchmarks
"We provide a database for evaluating computerized image-based prediction of the 7-point skin lesion malignancy checklist.
1 paper · 0 benchmarks
A Ball-Collision Dataset (ABCD) serves as a comprehensive benchmark for investigating the interaction dynamics of moving objects within 3D environments.
1 paper · 1 benchmark
This dataset is part of the publication "A bi-atrial statistical shape model for large-scale in silico studies of human atria: Model development and application to ECG simulations" by Nagel et al.
1 paper · 0 benchmarks
This is a dataset with curb annotations by using 3D LiDAR data and we build this dataset based on the SemanticKITTI dataset.
1 paper · 0 benchmarks
This dataset is meant to be used to develop models for next-day fire hazard forecasting in Greece.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A curated evaluation dataset for end-to-end Relation Extraction of relationships between organisms and natural-products.
1 paper · 0 benchmarks
We present a dataset of dialogs in which journalists of The Guardian replied to reader comments and identify the reasons why.
1 paper · 0 benchmarks
The dataset contains aerial agricultural images of a potato field with manual labels of healthy and stressed plant regions.
1 paper · 1 benchmark
This is a dataset of 583,437 tweets by 155,715 users that were censored between 2012-2020 July.
1 paper · 0 benchmarks
Using the CZI Software Mentions Dataset and ecosyste.ms we create a graph of papers, their mentioned software, and recursive dependencies of each piece of software across 466,000 papers and three software registries (PyPI, CRAN, and…
1 paper · 0 benchmarks
This dataset contains 9 different seafood types collected from a supermarket in Izmir, Turkey for a university-industry collaboration project at Izmir University of Economics, and this work was published in ASYU 2020.
1 paper · 0 benchmarks
This dataset contains data of 125 1-hour simulations of ship motion during various sea states performing random maneuvers in 4 degrees of freedom (surge-sway-yaw-roll).
1 paper · 0 benchmarks
This dataset contains 304 manual evaluations of class-level software maintainability, drawn from 5 open-source projects: ArgoUML, Art of Illusion, Diary Management, JUnit 4, JSweet.
1 paper · 0 benchmarks
This dataset contains a collection of papers retrieved by using a PRISMA systematic review of Open Data and Public Domain data in Agriculture.
1 paper · 0 benchmarks
A View From Somewhere (AVFS)—a dataset of 638,180 face similarity judgments over 4,921 faces.
1 paper · 0 benchmarks
This dataset contains a collection of 131 X-ray CT scans of pieces of modeling clay (Play-Doh) with various numbers of stones inserted, retrieved in the FleX-ray lab at CWI.
1 paper · 0 benchmarks
This dataset is a collection of undirected and unweighted LFR benchmark graphs as proposed by Lancichinetti et al.
1 paper · 0 benchmarks
This dataset contains a collection of 235800 X-ray projections of 131 pieces of modeling clay (Play-Doh) with various numbers of stones inserted.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Neonatal seizures are a common emergency in the neonatal intensive care unit (NICU).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is a ground truth dataset that is used to identify bots.
1 paper · 0 benchmarks
Dataset for A probabilistic forecast methodology for volatile electricity prices in the Australian National Electricity Market
1 paper · 0 benchmarks
The dataset contains the following data from successful and failed executions of the Toyota HSR robot placing a book on a shelf.
1 paper · 0 benchmarks
A tutorial for the improved Loewner Framework for modal analysis
1 paper · 0 benchmarks
A-AVA (Actor-identified Spatiotemporal Action Detection)
A new Actor-identified A-AVA dataset based on the existing AVA dataset and the TAO dataset, by assigning the unique actor identity and actions to each actor.
1 paper · 0 benchmarks
This dataset is based on FB15k237 and a pre-trained language-model-based KGE.
1 paper · 0 benchmarks
This dataset is based on WN18RR and a pre-trained language-model-based KGE.
1 paper · 0 benchmarks
A2Dre (Subset of A2D Sentences which are not trivial)
We obtain A2Dre by selecting only instances that were labeled as non-trivial, which are 433 REs from 190 videos.
1 paper · 1 benchmark
A2Dre+ (Extension of A2D sentences where trivial cases where filtered)
A2Dre is a subset from the A2D test set including $433$~\textit{non-trivial} REs.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.