Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 170 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8113–8160 of 12,172
This dataset contains recordings of 32 sound producing insect species with a total 335 files and a length of 57 minutes.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
~1M Flickr images from the XX century-aged from the 1910s to 1990s.
1 paper · 0 benchmarks
A dataset of stance-labeled GW sentences.
1 paper · 0 benchmarks
The laparoscopic surgery dataset is associated with our International Journal of Computer Assisted Radiology and Surgery (IJCARS) publication titled “DeSmoke-LAP: Improved Unpaired Image-to-Image Translation for Desmoking in Laparoscopic…
1 paper · 0 benchmarks
DeVAn (Dense Video Annotation for Video-Language Models)
DeVAn is a multi-modal dataset containing 8.5K video clips carefully selected from previously published YouTube-based video datasets (YouTube-8M and YT-Temporal-1B) that integrate visual and auditory information.
1 paper · 0 benchmarks
Decaf (deformation capture of faces interacting with hands)
We introduce the first monocular motion capture method from a video that regresses 3D hand and face motions along with deformations arising from their interactions.
1 paper · 0 benchmarks
DeePore (Deep learning for rapid characterization of porous materials)
DeePore is a deep learning workflow for rapid estimation of a wide range of porous material properties based on the binarized micro–tomography images.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We present a major update upon the original version of Deep Fashion 3D dataset .
1 paper · 0 benchmarks
This dataset inclue multi-spectral acquisition of vegetation for the conception of new DeepIndices.
1 paper · 1 benchmark
Release for uploading scripts and data to Zenodo Deep Neural Network Training Script incl.
1 paper · 0 benchmarks
The dataset contains two Pareto-fronts: - The Pareto-front for the 2-objective problem - The Pareto-front for the 3-objective problem Each Pareto-front contains a set of points, with coordinates given by their objectives.
1 paper · 0 benchmarks
Deep Soccer Captioning is a dataset consists of 22k caption-clip pairs and three visual features (images, optical flow, inpainting) for 500 hours of SoccerNet videos.
1 paper · 0 benchmarks
The Deep Thermal Imaging dataset consists of two main datasets: - DeepTherm I (Indoor materials) - 15 indoor materials were used to create the dataset DeepTherm I which consists of 14,860 processed thermal images (average count of data for…
1 paper · 0 benchmarks
There are 537 RGB jpg images of cracks and corresponding png binary segmentation masks of crack: a training set with 300 images and a testing set with 237 images.
1 paper · 0 benchmarks
DeepFigures Open (Extracting Scientific Figures with Distantly Supervised Neural Networks)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
DeepGraviLens is a data set of simulated gravitational lenses consisting of images associated with brightness variation time series.
1 paper · 0 benchmarks
DeepLocCross is a localization dataset that contains RGB-D stereo images captured at 1280 x 720 pixels at a rate of 20 Hz.
1 paper · 0 benchmarks
The code and database provided in this repository are related to the paper "DeepNetBeam: A Framework for the Analysis of Functionally Graded Porous Beams", which explores the application of various machine-learning techniques for the…
1 paper · 0 benchmarks
DeepParliament is a legal domain Benchmark Dataset that gathers bill documents and metadata and performs various bill status classification tasks.
1 paper · 0 benchmarks
Two versions of the dataset are offered: one is the full dataset used to train the models in DeformPAM, and the other is a mini dataset for easier examination.
1 paper · 0 benchmarks
For each DLO, we collect 350 seconds of dynamic trajectory data in the real-world using the motion capture system at a frequency of 100 Hz.
1 paper · 0 benchmarks
Delaunay triangulation dataset for 5, 10, 15, 20 points.
1 paper · 0 benchmarks
Demande Dataset contains the features and probabilites of ten different functions.
1 paper · 0 benchmarks
Demonstration video of the Stickbug Robot
1 paper · 0 benchmarks
This is the data regarding the pre-generated demonstration and experience replay for the proposed Deep-GRAIL algorithm.
1 paper · 0 benchmarks
Corpus for argument mining in legal documents, composed of 40 decisions of the Court of Justice of the European Union on matters of fiscal state aid
1 paper · 0 benchmarks
Source: Single-cell RNA-seq reveals dynamic, random monoallelic gene expression in mammalian cells
1 paper · 1 benchmark
Dense Forest Trail is an UAV dataset collected from a variety of simulated environment in Unreal Engine.
1 paper · 0 benchmarks
DensePose-Track is a dataset of videos where selected frames are annotated in the traditional DensePose manner.
1 paper · 0 benchmarks
Depth VIDIT (Virtual Image Dataset for Illumination Transfer)
VIDIT is a reference evaluation benchmark and to push forward the development of illumination manipulation methods.
1 paper · 0 benchmarks
Provide: 10 pickle files 8 pickle files are used to generate depth maps 2 pickle files are data of fronto parallel texture with their ground truth depth Each pickle file contains the parameter of the camera system (aperture size, optical…
1 paper · 0 benchmarks
A dataset of 100K synthetic images of skin lesions, ground-truth (GT) segmentations of lesions and healthy skin, GT segmentations of seven body parts (head, torso, hips, legs, feet, arms and hands), and GT binary masks of non-skin regions…
1 paper · 0 benchmarks
Desert Locus is a animal pose estimation dataset for desert locuses.
1 paper · 1 benchmark
A curated dataset of 221 question-answer-rationale triples capturing visualization design decisions and the reasoning behind them, derived from real-world student-authored narratives.
1 paper · 0 benchmarks
The code that created this dataset can be seen in https://github.com/nitzanfarhi/SecurityPatchDetection and can be reproduced by running: console python datacollection\createdataset.py --all -o datacollection\data Notice that this dataset…
1 paper · 0 benchmarks
Inspired by OpenAI dexterous in-hand manipulation, we collected a synthetic RGB-D dataset of a Shadow Hand robot manipulating a cube towards arbitrary goal configurations.
1 paper · 0 benchmarks
Dhoroni (Dhoroni: A Multi-Perspective Bengali Climate Change and Environmental News Dataset)
Climate change poses critical challenges globally, disproportionately affecting low-income countries that often lack resources and linguistic representation on the international stage.
1 paper · 1 benchmark
DiaKG is a high-quality Chinese dataset for Diabetes knowledge graph.
1 paper · 0 benchmarks
DiaSafety is a comprehensive dialogue safety dataset.
1 paper · 0 benchmarks
Contains Diabetic Foot Ulcers (DFU) from different patients.
1 paper · 0 benchmarks
This dataset contains anonymized data from patients seen at the Hospital Israelita Albert Einstein, at São Paulo, Brazil, and who had samples collected to perform the SARS-CoV-2 RT-PCR and additional laboratory tests during a visit to the…
1 paper · 0 benchmarks
Dialog-based Language Learning dataset is designed to measure how well models can perform at learning as a student given a teacher’s textual responses to the student’s answer (as well as potentially receiving an external real-valued reward…
1 paper · 0 benchmarks
DialogCC is a large-scale multi-modal dialogue dataset, which covers diverse real-world topics and various images per dialogue.
1 paper · 0 benchmarks
This dataset curates quantitative transparency disclosures about the online sexual exploitation of minors.
1 paper · 0 benchmarks
DigiCall (DigiCall: Earning Calls Dataset)
We release 3.691 earning call transcripts and also annotated data set, labeled particularly for the digital strategy maturity by linguists.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.