Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 129 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6145–6192 of 12,172
This dataset captures incompressible fluid dynamics around a 2D circular cylinder within a channel: ∇·𝐮 = 0, 𝐮 + (𝐮 ·∇) 𝐮 = ν∇² 𝐮 - 1/ρ ∇p, with boundary conditions set for velocity and pressure.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
CytoImageNet (CytoImageNet: A large-scale pretraining dataset for bioimage transfer learning)
CytoImageNet is a large-scale pretraining dataset of microscopy images (890K, 894 classes).
2 papers · 0 benchmarks
Czech restaurant information is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language.
2 papers · 1 benchmark
DACCORD is a new dataset dedicated to the task of automatically detecting contradictions between sentences in French.
2 papers · 0 benchmarks
The DAPlankton dataset consists of over 110k expert-labeled plankton images.
2 papers · 0 benchmarks
DARK FACE (DARK FACE: Face Detection in Low Light Condition)
DARK FACE dataset provides 6,000 real-world low light images captured during the nighttime, at teaching buildings, streets, bridges, overpasses, parks etc., all labeled with bounding boxes for of human face, as the main training and/or…
2 papers · 0 benchmarks
This is an SDQC stance-annotated Reddit dataset for the Danish language generated within a thesis project.
2 papers · 0 benchmarks
A large-scale multi-scene dataset for stereo deblurring, containing 20,637 blurry-sharp stereo image pairs from 135 diverse sequences and their corresponding bidirectional disparities.
2 papers · 0 benchmarks
DBATES (DataBase of Audio features, Text and visual Expressions in competitive debate Speeches)
DBATES is a database of multimodal communication features extracted from debate speeches in the 2019 North American Universities Debate Championships (NAUDC).
2 papers · 0 benchmarks
DBRD (Dutch Book Reviews Dataset)
The DBRD (pronounced dee-bird) dataset contains over 110k book reviews along with associated binary sentiment polarity labels.
2 papers · 1 benchmark
The DCASE 2017 rare sound events dataset contains isolated sound events for three classes: 148 crying babies (mean duration 2.25s), 139 glasses breaking (mean duration 1.16s), and 187 gun shots (mean duration 1.32s).
2 papers · 0 benchmarks
Based on the DDD17 dataset, we select some image-event pairs to evaluate the segmentation performance, namely DDD17-SEG, which only serves as a test set.
2 papers · 1 benchmark
DEplain-APA-sent: A German Parallel Corpus for Sentence Simplification on News Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
2 papers · 1 benchmark
DEplain-web-sent: A German Parallel Corpus for Sentence Simplification on Web Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
2 papers · 1 benchmark
Danish Fungi 2020 (DF20) is a novel fine-grained dataset and benchmark.
2 papers · 1 benchmark
Object Detection data set created from the engine DeepGTAV, which is based on the video game GTAV.
2 papers · 0 benchmarks
Object Detection data set created from the engine DeepGTAV, which is based on the video game GTAV.
2 papers · 0 benchmarks
DGraphFin dataset which is pre-processed in TGN Style.
2 papers · 0 benchmarks
DHP19 (Dynamic Vision Sensor 3D Human Pose Dataset)
DHP19 is the first human pose dataset with data collected from DVS event cameras.
2 papers · 1 benchmark
DIALOCONAN is a dataset comprising over 3000 fictitious multi-turn dialogues between a hater and an NGO operator, covering 6 targets of hate.
2 papers · 0 benchmarks
DIB-10K (DongNiao International Birds 10000)
Is a challenging image dataset which has more than 10 thousand different types of birds.
2 papers · 1 benchmark
The dataset contains digital ink drawings of diagrams with dynamic drawing information.
2 papers · 0 benchmarks
DISC-Law-SFT comprises two subsets, DISC-Law-SFT-Pair and DISC-Law-SFT-Triplet.
2 papers · 0 benchmarks
DISL (Fueling Research with A Large Dataset of Solidity Smart Contracts)
DISL The full dataset report is available at: https://arxiv.org/abs/2403.16861 The DISL dataset features a collection of 514, 506 unique Solidity files that have been deployed to Ethereum mainnet.
2 papers · 0 benchmarks
DL3DV-10K is a dataset of real-world videos with scene annotations and camera parameters.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
This is the dataset for the CGF 2021 paper "DONeRF: Towards Real-Time Rendering of Compact Neural Radiance Fields using Depth Oracle Networks".
2 papers · 1 benchmark
This is an open-source image captions dataset for the aesthetic evaluation of images.
2 papers · 0 benchmarks
DPPIN is a collection of dynamic networks, which consists of twelve generated dynamic protein-protein interaction networks of yeast cells, stored in twelve folders.
2 papers · 0 benchmarks
DRACO20K dataset is used for evaluating object canonicalization on methods that estimate a canonical frame from a monocular input image.
2 papers · 0 benchmarks
DRI Corpus (Dr. Inventor Multi-layer Scientific Corpus)
The Dr.
2 papers · 2 benchmarks
DSBI (Double-Sided Braille Image)
The Double-Sided Braille Image dataset (DSBI) is a large-scale dataset for Braille image recognition.
2 papers · 0 benchmarks
The dsd100 is a dataset of 100 full lengths of music tracks of different styles along with their isolated drums, bass, vocals, and other stems.
2 papers · 1 benchmark
Based on the DSEC dataset, we select some image-event pairs to evaluate the segmentation performance, namely DSEC-SEG, which only serves as a test set.
2 papers · 1 benchmark
In this paper, we introduce a novel benchmarking framework designed specifically for evaluations of data science agents.
2 papers · 0 benchmarks
In this paper, we introduce a novel benchmarking framework designed specifically for evaluations of data science agents.
2 papers · 0 benchmarks
In this paper, we introduce a novel benchmarking framework designed specifically for evaluations of data science agents.
2 papers · 0 benchmarks
DSIOD (Driving Scenario Input Output Dataset)
This dataset contains data which enables the evaluation of metamodels and approches for targeted test case selection without setting up test environments or performing test runs.
2 papers · 0 benchmarks
DTGB (Dynamic Text-attributed Graph Benchmark)
We introduce Dynamic Text-attributed Graph Benchmark (DTGB), a collection of large-scale, time-evolving graphs from diverse domains, with nodes and edges enriched by dynamically changing text attributes and categories.
2 papers · 0 benchmarks
Digital-Twin Tracking Dataset (DTTD) is a novel RGB-D dataset to enable further research of the problem and extend potential solutions towards longer ranges and mm localization accuracy.
2 papers · 0 benchmarks
DUC 2007 (Document Understanding Conferences)
There is currently much interest and activity aimed at building powerful multi-purpose information systems.
2 papers · 0 benchmarks
DaNewsroom (DaNewsroom: A Large-scale Danish Summarisation Dataset)
The first large-scale non-English language dataset specifically curated for automatic summarisation.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Darpa OpTC (Darpa Operationally Transparent Cyber (OpTC) Dataset)
Operationally Transparent Cyber (OpTC) was a technology transition pilot study funded under Boston Fusion Corp.'s Cyber APT Scenarios for Enterprise Systems (CASES) project.
2 papers · 0 benchmarks
This experiment was performed in order to empirically measure the energy use of small, electric Unmanned Aerial Vehicles (UAVs).
2 papers · 1 benchmark
This repository contains the database of the FEM simulation of axially impacted various configurations of the square crash boxes.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.