Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 112 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5329–5376 of 12,172
EMOVO (EMOVO Corpus: an Italian emotional speech database)
This article describes the first emotional corpus, named EMOVO, applicable to Italian language,.
3 papers · 0 benchmarks
ENST Drums (ENST-Drums: an extensive audio-visual database for drum signals processing)
ENST-Drums: an extensive audio-visual database for drum signals processing Olivier Gillet and Gaël Richard GET / ENST, CNRS LTCI, 37 rue Dareau, 75014 Paris, France The ENST-Drums database is a large and varied research database for…
3 papers · 0 benchmarks
From Grounded Human-Object Interaction Hotspots from Video (ICCV'19): We collect annotations for interaction keypoints on EPIC Kitchens in order to quantitatively evaluate our method in parallel to the OPRA dataset (where annotations are…
3 papers · 1 benchmark
ERA5 (The 5th generation of ECMWF reanalysis data)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
ES-ImageNet is a large-scale event-stream dataset for SNNs and neuromorphic vision.
3 papers · 0 benchmarks
The BIWI Walking Pedestrians dataset consists of walking pedestrians in busy scenarios from a birds eye view.
3 papers · 0 benchmarks
A massive, deduplicated corpus of 7.4M Python files from GitHub.
3 papers · 0 benchmarks
The EXEQ-300k dataset contains 290,479 detailed questions with corresponding math headlines from Mathematics Stack Exchange.
3 papers · 0 benchmarks
EasyPortrait (Face Parsing and Portrait Segmentation Dataset)
We introduce a large-scale image dataset EasyPortrait for portrait segmentation and face parsing.
3 papers · 0 benchmarks
Ego4D-HCap is a hierarchical video captioning dataset comprised of a three-tier hierarchy of captions: short clip-level captions, medium-length video segment descriptions, and long-range video-level summaries.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
ElBa (ElBa: Element Based Textures Dataset)
ElBa is composed of procedurally-generated realistic renderings, where we vary in a continuous way element shapes and colors and their distribution, to generate 30K texture images with different local symmetry, stationarity, and density of…
3 papers · 0 benchmarks
Email Thread Summarization (EmailSum) is a dataset which contains human-annotated short (<30 words) and long (<100 words) summaries of 2,549 email threads (each containing 3 to 10 emails) over a wide variety of topics.
3 papers · 2 benchmarks
Emotional Dialogue Acts data contains dialogue act labels for existing emotion multi-modal conversational datasets.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 1 benchmark
The original dataset from the reference consists of 5 different folders, each with 100 files, with each file representing a single subject/person.
3 papers · 1 benchmark
ErAConD (Error Annotated Conversational Dialog Dataset for Grammatical Error Correction)
ErAConD is a novel GEC dataset consisting of parallel original and corrected utterances drawn from open-domain chatbot conversations.
3 papers · 0 benchmarks
EuroCrops is a dataset for automatic vegetation classification from multi-spectral and multi-temporal satellite data, annotated with official LIPS reporting data from countries of the European Union, curated by the Technical University of…
3 papers · 0 benchmarks
EventNarrative is a knowledge graph-to-text dataset from publicly available open-world knowledge graphs.
3 papers · 1 benchmark
This is a medical multiple-choice dataset with explanations which can be used to interpret the answer.
3 papers · 0 benchmarks
ExtMarker (3D motion of chest external markers)
Three-dimensional position of external markers placed on the chest and abdomen of healthy individuals breathing during intervals from 73s to 222s.
3 papers · 1 benchmark
A new spatio-temporal benchmark dataset (Hurricane), is suited for forecasting during extreme events and anomalies.
3 papers · 1 benchmark
The FB15k-237-low dataset is a variation of the FB15k-237 dataset where relations with a low number of triplets are kept.
3 papers · 0 benchmarks
FCDB (Fashion Culture DataBase)
Consists of 76 million geo-tagged images in 16 cosmopolitan cities.
3 papers · 0 benchmarks
FEAFA+ is a dataset for Facial expression analysis and 3D Facial animation.
3 papers · 0 benchmarks
FFHQ-Text is a small-scale face image dataset with large-scale facial attributes, designed for text-to-face generation & manipulation, text-guided facial image manipulation, and other vision-related tasks.
3 papers · 0 benchmarks
Sharan, Lavanya, Ruth Rosenholtz, and Edward Adelson.
3 papers · 1 benchmark
FOBIE (Focused Open Biological Information Extraction)
The Focused Open Biology Information Extraction (FOBIE) dataset aims to support IE from Computer-Aided Biomimetics.
3 papers · 0 benchmarks
FOD in Airports (FOD-A) is an image dataset of FOD, Foreign Object Degris, which consists of 31 object categories and over 30,000 annotation instances.
3 papers · 0 benchmarks
FREDo is a Few-Shot Document-Level Relation Extraction Benchmark based on DocRED and SciERC.
3 papers · 2 benchmarks
A Few-Shot Learning Dataset of Molecules.
3 papers · 0 benchmarks
FSDD (Free Spoken Digit Dataset)
Free Spoken Digit Dataset (FSDD) is a simple audio/speech dataset consisting of recordings of spoken digits in wav files at 8kHz.
3 papers · 0 benchmarks
Arabic handwriting dataset.
3 papers · 1 benchmark
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
FakeNewsAMT & Celebrity include two novel datasets for the task of fake news detection, covering seven different news domains.
3 papers · 0 benchmarks
This is a dataset for segmentation and classification of epistemic activities in diagnostic reasoning texts.
3 papers · 0 benchmarks
FanOutQA is a high quality, multi-hop, multi-document benchmark for large language models using English Wikipedia as its knowledge base.
3 papers · 0 benchmarks
This dataset is derived from the Stack Overflow Data hosted by kaggle.com and available to query through Kernels using the BigQuery API: https://www.kaggle.com/stackoverflow/stackoverflow
3 papers · 0 benchmarks
Fetoscopic Placental Vessel Segmentation and Registration (FetReg) is a large-scale multi-centre dataset for the development of generalized and robust semantic segmentation and video mosaicking algorithms for the fetal environment with a…
3 papers · 0 benchmarks
FindingEmo is an image dataset containing annotations for 25k images, specifically tailored to Emotion Recognition.
3 papers · 0 benchmarks
FireRisk (FireRisk: A Remote Sensing Dataset for Fire Risk Assessment)
In this work, we propose a novel remote sensing dataset, FireRisk, consisting of 7 fire risk classes with a total of 91 872 labelled images for fire risk assessment.
3 papers · 1 benchmark
Fishnet Open Images Database is a large dataset of EM imagery for fish detection and fine-grained categorisation onboard commercial fishing vessels.
3 papers · 0 benchmarks
Fitness-AQA (Fitness Action Quality Assessment [ECCV 2022])
Largest, first-of-its-kind, in-the-wild, fine-grained workout/exercise posture analysis dataset, covering three different exercises: BackSquat, Barbell Row, and Overhead Press.
3 papers · 0 benchmarks
The Five-Billion-Pixels dataset contains more than 5 billion labeled pixels of 150 high-resolution Gaofen-2 (4 m) satellite images, annotated in a 24-category system covering artificial-constructed, agricultural, and natural classes.
3 papers · 0 benchmarks
FixMyPose is a dataset for automated pose correction.
3 papers · 0 benchmarks
FloDial (Flowchart Grounded Dialogs Dataset)
Flowchart Grounded Dialog Dataset (FloDial) is a corpus of troubleshooting dialogs between a user and an agent collected using Amazon Mechanical Turk.
3 papers · 0 benchmarks
Florence 4D is a dataset that consists of dynamic sequences of 3D face models, where a combination of synthetic and real identities exhibit an unprecedented variety of 4D facial expressions, with variations that include the classical…
3 papers · 0 benchmarks
GFP-GOWT1 mouse stem cells Dr.
3 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.