Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 100 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4753–4800 of 12,172
FRMT (Few-shot Region-aware Machine Translation)
FRMT is a dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation.
4 papers · 4 benchmarks
FacetSum is a faceted summarization dataset for scientific documents.
4 papers · 1 benchmark
The FakeMusicCaps dataset contains total of 27605 10 seconds music tracks corresponding to almost 77 hours, generated using 5 different Text-To-Music (TTM) models.
4 papers · 0 benchmarks
FewSOL (A Dataset for Few-Shot Object Learning in Robotic Environments)
The Few-Shot Object Learning (FewSOL) dataset can be used for object recognition with a few images per object.
4 papers · 0 benchmarks
FiNER-139 is comprised of 1.1M sentences annotated with eXtensive Business Reporting Language (XBRL) tags extracted from annual and quarterly reports of publicly-traded companies in the US.
4 papers · 0 benchmarks
The FieldSAFE dataset is a multi-modal dataset for obstacle detection in agriculture.
4 papers · 0 benchmarks
Pretrain: 200k Instruction: 100k
4 papers · 0 benchmarks
The first NER dataset in the field of traffic, which is to extract the characteristics and attributes of the vehicle on the road.
4 papers · 2 benchmarks
The dataset consists of 96 terrain-corrected (Level-1T) scenes from Landsat 8 OLI and TIRS, covering diverse biomes.
4 papers · 1 benchmark
Finer (Finnish News Corpus for Named Entity Recognition)
Finnish News Corpus for Named Entity Recognition (Finer) is a corpus that consists of 953 articles (193,742 word tokens) with six named entity classes (organization, location, person, product, event,and date).
4 papers · 0 benchmarks
Finnish Paraphrase Corpus is a fully manually annotated paraphrase corpus for Finnish containing 53,572 paraphrase pairs harvested from alternative subtitles and news headings.
4 papers · 0 benchmarks
FollowIR tests whether retrieval models can take queries that have fine-grained instructions in them
4 papers · 0 benchmarks
A dataset of high resolution, textured scans of articulated left feet, useful for 3D shape representation learning.
4 papers · 0 benchmarks
The Fraunhofer IPA Bin-Picking dataset is a large-scale dataset comprising both simulated and real-world scenes for various objects (potentially having symmetries) and is fully annotated with 6D poses.
4 papers · 0 benchmarks
Fruits 360 (A dataset of images containing fruits, vegetables, nuts and seeds)
Fruits-360 dataset: A dataset of images containing fruits, vegetables, nuts and seeds Version: 2025.03.24.0 Content The following fruits, vegetables and nuts and are included: Apples (different varieties: Crimson Snow, Golden, Golden-Red,…
4 papers · 0 benchmarks
GAMMA releases the world's first multi-modal dataset for glaucoma grading, which was provided by the Sun Yat-sen Ophthalmic Center of Sun Yat-sen University in Guangzhou, China.
4 papers · 0 benchmarks
GDSC (Genomics of Drug Sensitivity in Cancer)
We have characterized 1000 human cancer cell lines and screened them with 100s of compounds.
4 papers · 1 benchmark
GGPONC (German Guideline Program in Oncology NLP Corpus)
German Guideline Program in Oncology NLP Corpus (GGPONC) is a German language corpus based on clinical practice guidelines for oncology.
4 papers · 0 benchmarks
Grammatical error correction dataset for text from Wikipedia.
4 papers · 0 benchmarks
GVLM (Global Very-High-Resolution Landslide Mapping)
For change detection tasks, current open-source datasets mainly focus on building extraction (e.g., WHU building dataset and LEVIR-CD dataset) (Chen and Shi, 2020; Ji et al., 2018) and urban development monitoring (e.g., SECOND dataset,…
4 papers · 1 benchmark
Gazeta is a dataset for automatic summarization of Russian news.
4 papers · 1 benchmark
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
Glass (Glass Identification)
From USA Forensic Science Service; 6 types of glass; defined in terms of their oxide content (i.e.
4 papers · 2 benchmarks
Goal is a novel dataset of football (or 'soccer') highlights videos with transcribed live commentaries in English.
4 papers · 0 benchmarks
GoodSounds dataset contains around 28 hours of recordings of single notes and scales played by 15 different professional musicians, all of them holding a music degree and having some expertise in teaching.
4 papers · 0 benchmarks
What do doctors do when a patient has trouble breathing?
4 papers · 0 benchmarks
H-DIBCO 2010 is the International Document Image Binarization Contest which is dedicated to handwritten document images organized in conjunction with ICFHR 2010 conference.
4 papers · 0 benchmarks
HC-STVG1 (Human-centric Spatio-Temporal Video Grounding)
The newly proposed HC-STVG task aims to localize the target person spatio-temporally in an untrimmed video.
4 papers · 1 benchmark
HKR (Handwritten Kazakh and Russian (HKR) Database for Text Recognition)
The database is written in Cyrillic and shares the same 33 characters.
4 papers · 1 benchmark
Models character profiles and gives dialogue agents the ability to learn characters' language styles through their HLAs.
4 papers · 0 benchmarks
HOMER (Household Object Movements from Everyday Routines)
The Household Object Movements from Everyday Routines (HOMER) dataset is composed of routine behaviors for five households, spanning 50 days for the train split and 10 days for test split.
4 papers · 0 benchmarks
HPD (Head-Pose Detection)
These images were generated using Blender and IEE-Simulator with different head-poses, where the images are labelled according to nine classes (straight, turned bottom-left, turned left, turned top-left, turned bottom-right, turned right,…
4 papers · 0 benchmarks
The data simulate simple consumer banking interactions, containing about 23 hours of audio from 1,446 human-human conversations between 59 unique speakers.
4 papers · 0 benchmarks
Hate speech has become one of the most significant issues in modern society, with implications in both the online and offline worlds.
4 papers · 1 benchmark
We introduce an object detection dataset in challenging adverse weather conditions covering 12000 samples in real-world driving scenes and 1500 samples in controlled weather conditions within a fog chamber.
4 papers · 1 benchmark
Analyzing the surgical workflow is a prerequisite for many applications in computer assisted surgery (CAS), such as context-aware visualization of navigation information, specifying the most probable tool required next by the surgeon or…
4 papers · 2 benchmarks
Images with paired ground-truth caption hierarchies
4 papers · 0 benchmarks
This dataset contains five notable histological artifacts: blur, blood (hemorrhage), air bubbles, folded tissue, and damaged tissue.
4 papers · 1 benchmark
The historical color image dataset is collected for the task of automatically estimating the age of historical color photos.
4 papers · 0 benchmarks
Houses3K is a dataset of 3000 textured 3D house models.
4 papers · 0 benchmarks
HowTo100M Adverbs is a subset from HowTo100M with mined adverbs from 83 tasks in HowTo100M.
4 papers · 1 benchmark
HuPR (Human Pose with Millimeter Wave Radar)
HuPR is a human pose estimation benchmark is created using cross-calibrated mmWave radar sensors and a monocular RGB camera for cross-modality training of radar-based human pose estimation.
4 papers · 0 benchmarks
Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains.
4 papers · 0 benchmarks
Human-M3 is an outdoor multi-modal multi-view multi-person human pose database which includes not only multi-view RGB videos of outdoor scenes but also corresponding pointclouds.
4 papers · 0 benchmarks
Propose a dataset which adopts multi-channel visual input.
4 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
Money laundering is a multi-billion dollar issue.
4 papers · 0 benchmarks
A dataset for inter-personal relationship extraction which aims to facilitate information extraction and knowledge graph construction research.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.