Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 107 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5089–5136 of 12,172
WikiNLDB is a novel dataset for training Natural Language Databases (NLDBs) which is generated by transforming structured data from Wikidata into natural language facts and queries.
4 papers · 0 benchmarks
The WikiSem500 dataset contains around 500 per-language cluster groups for English, Spanish, German, Chinese, and Japanese (a total of 13,314 test cases).
4 papers · 0 benchmarks
Wikipedia Citations is a comprehensive dataset of citations extracted from Wikipedia.
4 papers · 0 benchmarks
Wikipedia Generation is a dataset for article generation from Wikipedia from references at the end of Wikipedia page and the top 10 search results for the Wikipedia topic.
4 papers · 0 benchmarks
The WikipediaGS dataset was created by extracting Wikipedia tables from Wikipedia pages.
4 papers · 2 benchmarks
The Wino-X dataset is a multilingual collection of Winograd Schemas.
4 papers · 0 benchmarks
The World Mortality Dataset contains weekly, monthly, or quarterly all-cause mortality data from 103 countries and territories.
4 papers · 0 benchmarks
The PAirMax dataset is a collection of images for evaluating the performance of pansharpening algorithms.
4 papers · 1 benchmark
SRL is the task of extracting semantic predicate-argument structures from sentences.
4 papers · 0 benchmarks
XED is a multilingual fine-grained emotion dataset.
4 papers · 0 benchmarks
XLING (XLING BLI Dataset)
The XLING BLI Dataset contains bilingual dictionaries for 28 language pairs.
4 papers · 0 benchmarks
A large-scale dataset built on questions from TyDi QA lacking same-language answers.
4 papers · 0 benchmarks
XTD10 is a dataset for cross-lingual image retrieval and tagging consisting of the MSCOCO2014 caption test dataset annotated in 7 languages that were collected using a crowdsourcing platform.
4 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
Youku-mPLUG is a large Chinese high-quality video-language dataset which is collected from Youku.com, a well-known Chinese video-sharing website, with strict criteria of safety, diversity, and quality.
4 papers · 0 benchmarks
ZhihuRec dataset is collected from a knowledge-sharing platform (Zhihu), which is composed of around 100M interactions collected within 10 days, 798K users, 165K questions, 554K answers, 240K authors, 70K topics, and more than 501K user…
4 papers · 0 benchmarks
aethel (Automatically Extracted Theorems from Lassy)
A dataset of approximately 75,000 phrases and sentences, syntactically analyzed as typelogical derivations (i.e.
4 papers · 0 benchmarks
For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of physics to the many subdisciplines of computer science to everything in between, including…
4 papers · 1 benchmark
dMelodies is dataset of simple 2-bar melodies generated using 9 independent latent factors of variation where each data point represents a unique melody based on the following constraints: - Each melody will correspond to a unique scale…
4 papers · 0 benchmarks
The eBDtheque database is a selection of one hundred comic pages from America, Japan (manga) and Europe.
4 papers · 1 benchmark
eICU-CRD (eICU Collaborative Research Database)
The eICU Collaborative Research Database is a large multi-center critical care database made available by Philips Healthcare in partnership with the MIT Laboratory for Computational Physiology.
4 papers · 2 benchmarks
The eSports Sensors dataset contains sensor data collected from 10 players in 22 matches in League of Legends.
4 papers · 2 benchmarks
The dataset is described here: https://grobid-quantities.readthedocs.io/en/latest/guidelines.html
4 papers · 0 benchmarks
In response to the evolving challenges posed by small unmanned aerial vehicles (UAVs), which possess the potential to transport harmful payloads or independently cause damage, we introduce MMAUD: a comprehensive Multi-Modal Anti-UAV…
4 papers · 0 benchmarks
This dataset contains 1304 de-identified longitudinal medical records describing 296 patients.
4 papers · 1 benchmark
iShape is an irregular shape dataset for instance segmentation.
4 papers · 1 benchmark
This is a dataset for disentangling conversations on IRC, which is the task of identifying separate conversations in a single stream of messages.
4 papers · 3 benchmarks
jazznet is a dataset of piano patterns for music audio machine learning research.
4 papers · 0 benchmarks
Pn-summary is a dataset for Persian abstractive text summarization.
4 papers · 0 benchmarks
A dataset of more than ten thousand 3D scans of real objects.
4 papers · 0 benchmarks
t4d (Thinking is for Doing)
This dataset was generated by the code implementation found here: https://github.com/sachith-gunasekara/t4d
4 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
2devs is a publicly available dataset of fine-grained untangled code changes collected by recording the development sessions of two developers over the course of four months, and the corresponding manual clustering.
3 papers · 0 benchmarks
The dataset is designed specifically to solve a range of computer vision problems (2D-3D tracking, posture) faced by biologists while designing behavior studies with animals.
3 papers · 0 benchmarks
3DMAD (3D Mask Attack Dataset)
The 3D Mask Attack Database (3DMAD) is a biometric (face) spoofing database.
3 papers · 0 benchmarks
97 synthetic datasets consists of 97 datasets (as illustrated in the figure) and can be used to test graph-based clustering algorithms.
3 papers · 1 benchmark
Robot grasping is often formulated as a learning problem.
3 papers · 0 benchmarks
This dataset gathers 10,874 title and abstract pairs from the ACL Anthology Network (until 2016).
3 papers · 1 benchmark
ACL-Fig is a large-scale automatically annotated corpus consisting of 112,052 scientific figures extracted from 56K research papers in the ACL Anthology.
3 papers · 0 benchmarks
ADE-Affordance is a new dataset that builds upon ADE20k, which contains annotations enabling such rich visual reasoning.
3 papers · 0 benchmarks
Attention Deficit Hyperactivity Disorder (ADHD) affects at least 5-10% of school-age children and is associated with substantial lifelong impairment, with annual direct costs exceeding $36 billion/year in the US.
3 papers · 0 benchmarks
AFEW-VA (AFEW-VA Database for Valence and Arousal Estimation In-The-Wild)
The AFEW-VA databaset is a collection of highly accurate per-frame annotations levels of valence and arousal, along with per-frame annotations of 68 facial landmarks for 600 challenging video clips.
3 papers · 0 benchmarks
AIR-Act2Act is a human-human interaction dataset for teaching non-verbal social behaviors to robots.
3 papers · 0 benchmarks
The AND Dataset contains 13700 handwritten samples and 15 corresponding expert examined features for each sample.
3 papers · 1 benchmark
Extension test cases of APPS, as well as generated code.
3 papers · 0 benchmarks
ARCH2S (Dataset, Benchmark for Learning Exterior Architectural Structures from Point Clouds)
Precise segmentation of architectural structures provides detailed information about various building components, enhancing our understanding and interaction with our built environment.
3 papers · 1 benchmark
The AROT-COV23 (ARabic Original Tweets on COVID-19 as of 2023) dataset is a large-scale collection of original Arabic tweets related to COVID-19, spanning from January 2020 to January 2023, and the period for which we collected the data…
3 papers · 0 benchmarks
AS-V2 (The All-Seeing Dataset v2)
We propose a novel task, termed Relation Conversation (ReC), which unifies the formulation of text generation, object localization, and relation comprehension.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.