Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 174 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8305–8352 of 12,172
EnvBench is a comprehensive benchmark for automating environment setup - an important task in software engineering.
1 paper · 0 benchmarks
Benchmark to evaluate the capability of LMs to consolidate and recall information from multiple training documents.
1 paper · 0 benchmarks
This repository contains three graph datasets for the UE traffic assignment problem on Sioux-Falls, Eastern-Massachusetts and Anaheim networks in both dgl and pyg formats.
1 paper · 0 benchmarks
ErhuPT (Erhu Playing Technique Dataset)
This dataset is an audio dataset containing about 1500 audio clips recorded by multiple professional players.
1 paper · 0 benchmarks
Provide: a high-level explanation of the dataset characteristics explain motivations and summary of its content potential use cases of the dataset Collection of Error Grid data files.
1 paper · 0 benchmarks
Essays (Stream-of-consciousness Essays)
J.
1 paper · 1 benchmark
The sampled 2-hop subgraphs centered on ICO-wallet accounts on the Ethereum Interaction graph.
1 paper · 0 benchmarks
The sampled 2-hop subgraphs centered on Mining accounts on the Ethereum Interaction graph.
1 paper · 0 benchmarks
aThis dataset provides NFT ownership traces and detection of potential wash trading activities across several prominent NFT collections on the Ethereum blockchain.
1 paper · 0 benchmarks
This dataset consists of charge densities of individual snapshots from a molecular dynamics trajectory (DFT simulations?).
1 paper · 0 benchmarks
The Euro-PVI dataset contains trajectories of pedestrians and bicyclists, with dense interactions with the ego-vehicle.
1 paper · 0 benchmarks
EuroSAT-C is an open-source data set comprising algorithmically generated corruptions applied to the EuroSAT test set following the concept of ImageNet-C.
1 paper · 0 benchmarks
The ConcoDisco Corpus is an English-French parallel corpus with discourse relations (DRs) and discourse connectives (DCs) annotations.
1 paper · 0 benchmarks
Eurovision 2018 official votes dataset.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is the supplemental data for our paper on how to benchmark registrations of serial sections with ground truths.
1 paper · 0 benchmarks
Contains 1000 semantic queries and the corresponding English, German and Portuguese verbalizations for EventKG - an event-centric knowledge graph with more than 970 thousand events.
1 paper · 0 benchmarks
Event-Stream Dataset is a robotic grasping dataset with 91 objects.
1 paper · 0 benchmarks
A corpus designed in analogy to the well-established English ISEAR emotion dataset.
1 paper · 0 benchmarks
EventEA is an event-centric entity alignment dataset, harvested from EventKG, DBpedia and Wikidata.
1 paper · 0 benchmarks
This dataset contains the broadcast video streams of handball matches along with synchronized official positional data and human event annotations for 125min raw data in summary.
1 paper · 0 benchmarks
EviLOG (Evidential Lidar Occupancy Grid Mapping)
The dataset contains synthetic training, validation and test data for occupancy grid mapping from lidar point clouds.
1 paper · 0 benchmarks
Intermediate annotations from the FEVER dataset that describe original facts extracted from Wikipedia and the mutations that were applied, yielding the claims in FEVER.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A dataset of illusions generated by the AI model EIGen.
1 paper · 0 benchmarks
The SF100 corpus of classes is a statistically representative sample of 100 Java projects from SourceForge, which is a popular open source repository (more than 300,000 projects with more than two million registered users).
1 paper · 0 benchmarks
ExAIS (ExAIS_SMS Spam dataset)
This ExAISSMS Spam dataset was a project conducted at the Federal University of Agriculture, Abeokuta, Nigeria with the aim of building an indigenous SMS Spam corpus with African-English context.
1 paper · 0 benchmarks
ExBAN (ExBAN Corpus (Explanations for BAyesian Networks))
The ExBAN dataset: a corpus of NL explanations generated by crowd-sourced participants presented with the task of explaining simple Bayesian Network (BN) graphical representations.
1 paper · 0 benchmarks
ExLPose (Extremely Low-light human Pose dataset)
We study human pose estimation in extremely low-light images.
1 paper · 0 benchmarks
ExPUNations is a humor dataset with such extensive and fine-grained annotations specifically for puns.
1 paper · 0 benchmarks
The ExaASC dataset is a dataset for Target-based Stance Detection in the Arabic Language that contains different types of targets like persons, entities and events.
1 paper · 0 benchmarks
This is an example data set for a hypothetical electronic products supply network.
1 paper · 0 benchmarks
Dataset to be used with the https://github.com/MathBioCU/WSINDyCellCluster code
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
0.This is experiment data for the following article: @misc{liu2021topic, title={Topic Model Supervised by Understanding Map}, author={Gangli Liu}, year={2021}, eprint={2110.06043}, archivePrefix={arXiv}, primaryClass={cs.CL} } 1.
1 paper · 0 benchmarks
This package contains the raw data / logs (fetched from WandB) for the experiments of the following publication: O.
1 paper · 0 benchmarks
Data files with the information required to replicate all the experiments reported in the paper: Linares López, Carlos; Herman, Ian, 2024.
1 paper · 0 benchmarks
Contains materials used in the experiments, raw results, and analysis scripts.
1 paper · 0 benchmarks
This repository presents the dataset used in the PerfCam's original paper.
1 paper · 0 benchmarks
This dataset is being used to evaluate PerfSim accuracy and speed against a real deployment in a Kubernetes cluster based on sfc-stress workloads.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
An image dataset containing simplified representations of trains.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Expository-Prose-V1 is a collection of specially-curated corpora gathered from diverse sources, ranging from research papers (arXiv) to European Parliament proceedings (EuroParl).
1 paper · 0 benchmarks
Neural network model files and Madgraph event generator outputs used as inputs to the results presented in the paper "Learning to discover: expressive Gaussian mixture models for multi-dimensional simulation and parameter inference in the…
1 paper · 0 benchmarks
To overcome the need for a full installation of a reverse geocoder such as Nominatim, we provide the post-processed output of the reverse geocoding for the MP-16 dataset along with the validation set (YFCC-Val26k) which originally…
1 paper · 0 benchmarks
Minecraft Corpus dataset with builder utterance annotations
1 paper · 0 benchmarks
A dataset of abdominal CT studies in NifTi format from the open-source medical data repository Medical Decathlon was utilized.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.