Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 230 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10993–11040 of 12,172
XA Bin-Picking is a point-cloud dataset comprising both simulated and real-world scenes with three industrial parts.
1 paper · 0 benchmarks
The World Ocean Database (WOD) is world's largest collection of uniformly formatted, quality controlled, publicly available ocean profile data.
1 paper · 0 benchmarks
A balanced dataset of color names and RGB values for training classifiers.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
XMIDI is a comprehensive, large-scale symbolic music dataset that includes accurate emotion and genre labels, consisting of 108,023 MIDI files.
1 paper · 0 benchmarks
We provide a new data set XWikiRef for the task of Cross-lingual Multi-document Summarization.
1 paper · 0 benchmarks
Synthetic dataset intended for benchmarking disentanglement frameworks.
1 paper · 0 benchmarks
Xamarin Q&A consists of two datasets of questions and answers for studying the development of cross-platform mobile applications using the Xamarin framework.
1 paper · 0 benchmarks
XiaChuFang Recipe Corpus contains recipes are from 下厨房 (XiaChuFang), a popular Chinese recipe sharing website.
1 paper · 0 benchmarks
YADL (Yet Another Data LAke)
Files composing the YADL data lake, for the paper "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes (Experiment, Analysis & Benchmark Paper)" Archives provided here follow the notation used for the experiments, which is…
1 paper · 0 benchmarks
The YCB-Ev dataset contains synchronized RGB-D frames and event data that enables evaluating 6DoF object pose estimation algorithms using these modalities.
1 paper · 0 benchmarks
The scales of the data accessible through internet search engines can reach hundreds of millions, or even billions.
1 paper · 0 benchmarks
The YFCC100M Fine-Grained Geolocation dataset is a subset of 100 a set of 36,146 YFCC100M images that had Flickr tags that could be identified as corresponding to one of the labels in the iNaturalist 2017 dataset.
1 paper · 0 benchmarks
An instance segmentation dataset of yeast cells in microstructures.
1 paper · 0 benchmarks
YJMob100K (YJMob100K: City-Scale and Longitudinal Dataset of Anonymized Human Mobility Trajectories)
Modeling and predicting human mobility trajectories in urban areas is an essential task for various applications including transportation modeling, disaster management, and urban planning.
1 paper · 0 benchmarks
YM2413-MDB is an 80s FM video game music dataset with multi-label emotion annotations.
1 paper · 0 benchmarks
Source: Single-cell RNA-Seq profiling of human preimplantation embryos and embryonic stem cells
1 paper · 1 benchmark
Data for the paper entitled Quantifying yeast colony morphologies with feature engineering from time-lapse photography by A.
1 paper · 0 benchmarks
YesBut Dataset (https://yesbut-dataset.github.io) Understanding satire and humor is a challenging task for even current Vision-Language models.
1 paper · 0 benchmarks
YorkTag provides pairs of sharp/blurred images containing fiducial markers and is proposed to train and qualitatively and quantitatively evaluate our model.
1 paper · 0 benchmarks
YouTube-GDD (YouTube-GDD: A challenging gun detection dataset with rich contextual information)
YouTubeGun Detection Dataset is collected from 343 high-definition YouTube videos and contains 5000 well-chosen images, in which 16064 instances of gun and 9046 instances of person are annotated.
1 paper · 0 benchmarks
YouTube-Hands includes 240 videos which are annotated with hand trajectories.
1 paper · 1 benchmark
We redistribute a suite of datasets as part of the YourMT3 project.
1 paper · 0 benchmarks
YoutubeGraph-Dyn is an evolving graph dataset generated from YouTube real-world interactions.
1 paper · 0 benchmarks
YouwikiHow is a dataset for Weakly-Supervised temporal Article Grounding (WSAG).
1 paper · 0 benchmarks
A corpus for two endangered languages of the Zaza-Gorani language family: Zazaki and Gorani.
1 paper · 0 benchmarks
ZeroKBC is comprehensive benchmark that covers all scenarios of zero-shot Knowledge Base Completion (KBC) task.
1 paper · 0 benchmarks
The first and the one open dataset for Russian finger- spelling, contained 1,593 annotated phrases and over 37 thousand HD+ videos.
1 paper · 1 benchmark
The Humbug Zooinverse dataset is a dataset of mosquito audio recordings.
1 paper · 0 benchmarks
ZuBuD (Zurich Buildings Database)
The goal of the ZuBuD Image Database is to share image data sets with researcheres around the world.
1 paper · 0 benchmarks
A more balanced version of ZuBuD.
1 paper · 0 benchmarks
Dataset Summary The dataset used to train and evaluate TunesFormer is collected from two sources: The Session and ABCnotation.com.
1 paper · 0 benchmarks
adVFed (Tencent Federated Advertising CVR Dataset)
Natural Vertical Partitioned CVR Dataset for Vertical Federated Learning This Dataset repo provides 2 industrial CVR Dataset for VFL research.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ai4st SLR (Research on AI for Software Testing Research, 2020-2025)
To check the validity of the ai4st ontology, an adapted, lightweight systematic literature review (SLR) was conducted to analyse related research.
1 paper · 0 benchmarks
A large-scale traffic sign and traffic light dataset with accurate 3D positioning and temporally consistent 3D bounding boxes of traffic management objects from up to 200 meters away.
1 paper · 0 benchmarks
A large-scale training dataset suffering from the defocus spread effect (DSE) is synthesized by applying an α-matte boundary defocus model to the VOC 2012 dataset.
1 paper · 0 benchmarks
https://huggingface.co/datasets/mediabiasgroup/anno-lexical
1 paper · 0 benchmarks
This dataset provides a curated collection of approved drug Simplified Molecular Input Line Entry System (SMILES) strings and their associated protein sequences.
1 paper · 0 benchmarks
This is a second public release of the arXMLiv dataset generated by the KWARC research group.
1 paper · 0 benchmarks
This is a dataset of scientific documents derived from arXiv.
1 paper · 0 benchmarks
A newly proposed dataset for local citation recommendation, consisting of 3.2 million local citation sentences along with the title and the abstract of both the citing and the cited papers.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTe# Arabic Img2MD Dataset Summary The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts.
1 paper · 0 benchmarks
A collection of various NLP datasets in Assamese.
1 paper · 0 benchmarks
bSDD (buildingSMART Data Dictionary)
The buildingSMART Data Dictionary (bSDD) is an online service that hosts classifications and their properties, allowed values, units and translations.
1 paper · 0 benchmarks
This is a high-quality dataset of annotated posts sampled from social media posts and annotated for misogyny.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.