Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 131 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6241–6288 of 12,172
ESB (End-to-End Speech Benchmark)
ESB is a benchmark for evaluating the performance of a single automatic speech recognition (ASR) system across a broad set of speech datasets.
2 papers · 0 benchmarks
The paper used 500 scanned Electronic Theses and Dissertation cover pages (i.e., front pages).
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
VD4UAV is an altitude-sensitive benchmark dataset designed to evade vehicle detection in Unmanned Aerial Vehicle (UAV) imagery.
2 papers · 2 benchmarks
Deep learning use for quantitative image analysis is exponentially increasing.
2 papers · 1 benchmark
To automatically generate Python and assembly programs used for security exploits, we curated a large dataset for feeding NMT techniques.
2 papers · 0 benchmarks
EasyCall corpus is a dysarthric speech command dataset in Italian.
2 papers · 0 benchmarks
EchoCP is an echocardiography dataset in cTTE targeting PFO (Patent foramen ovale) diagnosis.
2 papers · 0 benchmarks
EconLogicQA is a benchmark designed to test the sequential reasoning skills of large language models (LLMs) in economics, business, and supply chain management.
2 papers · 1 benchmark
The Eedi dataset contains from two school years (September 2018 to May 2020) of students’ answers to mathematics questions from Eedi, a leading educational platform which millions of students interact with daily around the globe.
2 papers · 0 benchmarks
Egoshots is a 2-month Ego-vision Dataset with Autographer Wearable Camera annotated "for free" with transfer learning.
2 papers · 0 benchmarks
EgoTV dataset consists of (task description, video) pairs with positive on negative task verification labels.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
The dataset available for download on this webpage represents a 5x5x5µm section taken from the CA1 hippocampus region of the brain, corresponding to a 1065x2048x1536 volume.
2 papers · 1 benchmark
This data was collected by performing a breadth-first search on the user-product-review graph until termination, meaning that it is a fairly comprehensive collection of English-language product data.
2 papers · 1 benchmark
An open corpus of Scientific Research papers which has a representative sample from across scientific disciplines.
2 papers · 0 benchmarks
EmoCause is a dataset of annotated emotion cause words in emotional situations from the EmpatheticDialogues valid and test set.
2 papers · 1 benchmark
EmoPars is a dataset of 30,000 Persian Tweets labeled with Ekman’s six basic emotions (Anger, Fear, Happiness, Sadness, Hatred, and Wonder).
2 papers · 0 benchmarks
The EmoTag1200 dataset is a collection of resources for analyzing the emotion and sentiment of emojis as well as tweets written in English.
2 papers · 0 benchmarks
1000 songs has been selected from Free Music Archive (FMA).
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
A challenge that consists of three tasks, each targeting a different requirement for in-clinic use.
2 papers · 1 benchmark
The protein-ligand complexes of PDBBind v2020 preprocessed as described in the paper "EquiBind: Geometric Deep Learning for Drug Binding Structure Prediction" with associated code at https://github.com/HannesStark/EquiBind Contained are…
2 papers · 0 benchmarks
This repository contains essays written by high school Brazilian students.
2 papers · 0 benchmarks
The sampled 2-hop subgraphs centered on Exchange accounts on the Ethereum Interaction graph.
2 papers · 0 benchmarks
Ethics (per ethics) dataset is created to test the knowledge of the basic concepts of morality.
2 papers · 1 benchmark
A multilingual etymological database extracted from the Wiktionary (described in Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0)
2 papers · 0 benchmarks
This dataset consists of 3,710 flood images, annotated by domain experts regarding their relevance with respect to three tasks (determining the flooded area, inundation depth, water pollution).
2 papers · 0 benchmarks
Recently, Text-to-Image (T2I) generation models have achieved significant advancements.
2 papers · 0 benchmarks
Event-Human3.6m is a challenging dataset for event-based human pose estimation by simulating events from the RGB Human3.6m dataset.
2 papers · 0 benchmarks
The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments.
2 papers · 0 benchmarks
ExHVV is a novel dataset that offers natural language explanations of connotative roles for three types of entities -- heroes, villains, and victims, encompassing 4,680 entities present in 3K memes.
2 papers · 0 benchmarks
This is a dataset of code snippets in StackOverflow that have been used in Github repositories by extending and adapting them.
2 papers · 0 benchmarks
ExpMRC is a benchmark for the Explainability evaluation of Machine Reading Comprehension.
2 papers · 0 benchmarks
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Extended Agriculture-Vision dataset comprises two parts: 1.
2 papers · 0 benchmarks
EyeCar is a dataset of driving videos of vehicles involved in rear-end collisions paired with eye fixation data captured from human subjects.
2 papers · 0 benchmarks
EyeDentify, a dataset specifically designed for pupil diameter estimation based on webcam images.
2 papers · 0 benchmarks
The Freebase Annotations of TREC KBA 2014 Stream Corpus with Timestamps (FAKBAT) is an extension of the FAKBA1 dataset that contains entity age and entity timestamp.
2 papers · 0 benchmarks
A new fraud detection dataset FDCompCN for detecting financial statement fraud of companies in China.
2 papers · 1 benchmark
A 360-degree fisheye-like version of the popular FDDB face detection dataset.
2 papers · 0 benchmarks
FDH (Flickr Diverse Humans)
The Flickr Diverse Humans (FDH) dataset consists of 1.53M images of human figures from the YFCC100M dataset.
2 papers · 0 benchmarks
A large dataset of routinely acquired maternal-fetal screening ultrasound images collected from two different hospitals by several operators and ultrasound machines.
2 papers · 0 benchmarks
FFHNet (FFHNet Dexterous Grasping Dataset)
https://syncandshare.lrz.de/getlink/fi9EZb33KiSAJ5rLHAkhg7/ffhnet-data.zip
2 papers · 0 benchmarks
FFSC (Face Forgery in the Semantic Context)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
FIB (Factual Inconsistency Benchmark)
Factual Inconsistency Benchmark (FIB) is a benchmark that focuses on the task of summarization.
2 papers · 0 benchmarks
FIJO (French Insurance Job Offer dataset)
This dataset was collected as part of the multidisciplinary project Femmes face aux défis de la transformation numérique : une étude de cas dans le secteur des assurances (Women Facing the Challenges of Digital Transformation: A Case Study…
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.