Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 180 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8593–8640 of 12,172
Collection of news websites in low-resource languages.
1 paper · 0 benchmarks
StoryBooks for 174 unique languages.
1 paper · 0 benchmarks
A Brazilian Portuguese TTS dataset featuring a female voice recorded with high quality in a controlled environment, with neutral emotion and more than 20 hours of recordings.
1 paper · 0 benchmarks
A database containing high sampling rate recordings of a single speaker reading sentences in Brazilian Portuguese with neutral voice, along with the corresponding text corpus.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Description This Dataset contains review information on Google map (ratings, text, images, etc.), business metadata (address, geographical info, descriptions, category information, price, open hours, and MISC info), and links (relative…
1 paper · 0 benchmarks
This dataset was curated for Search Engine Optimization (SEO) analysis tasks, including categorization and spam detection.
1 paper · 0 benchmarks
Source: Heterogeneity in Oct4 and Sox2 Targets Biases Cell Fate in 4-Cell Mouse Embryos
1 paper · 1 benchmark
GovDocs is a corpus of nearly 1 million documents that are freely available for research and may be, to the best of the authors' knowledge, freely redistributed.
1 paper · 0 benchmarks
Dataset introduced by Xifeng Yan et al.
1 paper · 0 benchmarks
Dataset introduced by Xifeng Yan et al.
1 paper · 0 benchmarks
The dataset is described in the README file in the official repository of the paper
1 paper · 0 benchmarks
Robotic grasp dataset for multi-object multi-grasp evaluation with RGB-D data.
1 paper · 0 benchmarks
GraspClutter6D is a large-scale real-world dataset for robust object perception and robotic grasping in cluttered environments.
1 paper · 0 benchmarks
A large-scale grasp pose detection dataset with a unified evaluation system.
1 paper · 0 benchmarks
Greek Parliament Proceedings is a curated dataset of the Greek Parliament Proceedings that extends chronologically from 1989 up to 2020.
1 paper · 0 benchmarks
GroundCap is a novel grounded image captioning dataset derived from MovieNet, containing 52,350 movie frames with detailed grounded captions.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Group LAW (LAW | SuiteSparse Matrix Collection)
URL: https://sparse.tamu.edu/LAW Laboratory for Web Algorithmics (LAW), Università degli Studi di Milano http://law.di.unimi.it/index.php When using matrices in the LAW/ group in the collection, please follow the citation instructions at…
1 paper · 0 benchmarks
For each problem, we provide 4 variants of prompts: 1.
1 paper · 0 benchmarks
Diverse guitar-playing motions about 1 hour long, including: • 12 major scales, • chromatic scales, • diverse chords, • arpeggios, • strumming and picking, • bends, • sliding, • vibrato, • palm mute, • natural harmonics, • artificial…
1 paper · 0 benchmarks
Guitar-TECHS (Guitar Tones/Techniques, Excerpts & Chords Dataset)
Guitar-TECHS is a comprehensive dataset featuring a variety of guitar techniques, musical excerpts, chords, and scales.
1 paper · 0 benchmarks
This is a high-quality dataset consisting of 14.8M utterances in English, extracted from processed dialogues from publicly available online books.
1 paper · 0 benchmarks
Gutenberg Poem Dataset is used for the next verse prediction component.
1 paper · 0 benchmarks
A data set of hourly time phrases from 52,183 fictional books.
1 paper · 0 benchmarks
The Human-to-Human-or-Object Interaction Dataset (H2O) dataset is a dataset for Human-Object Interaction (HOI) detection.
1 paper · 0 benchmarks
A dataset for pose estimation of hand when interacting with object and severe occlusions.
1 paper · 0 benchmarks
HA-ViD (HA-ViD: A Human Assembly Video Dataset)
Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry.
1 paper · 0 benchmarks
HAC (Hybrid Adverse Conditions)
HAC is a dataset for learning and benchmarking arbitrary Hybrid Adverse Conditions restoration.
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
HALvest-Geometric is a subset of HALvest: an academic citation network with 238,397 disambiguated authors and 18,662,037 scholarly papers.
1 paper · 0 benchmarks
HAM (Human-annotated Mappings)
HAM is a dataset for molecular graph partitioning.
1 paper · 0 benchmarks
HANS (Heuristic Analysis for NLI Systems)
The HANS (Heuristic Analysis for NLI Systems) dataset which contains many examples where the heuristics fail.
1 paper · 1 benchmark
HARD (Hotel Arabic-Reviews Dataset)
The Hotel Arabic-Reviews Dataset (HARD) contains 93700 hotel reviews in Arabic language.
1 paper · 1 benchmark
We introduce a challenging benchmark of graduate-level problems in applied mathematics that was fully developed as part of a university class.
1 paper · 0 benchmarks
HARRISON dataset is a benchmark on hashtag recommendation for real world images in social networks.
1 paper · 0 benchmarks
HASCD (Human Activity Segmentation Challenge Dataset)
HASCD (Human Activity Segmentation Challenge Dataset) contains 250 annotated multivariate time series capturing 10.7 h of real-world human motion smartphone sensor data from 15 bachelor computer science students.
1 paper · 0 benchmarks
HATIE (Human-Aligned benchmark for Text-guided Image Editing)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
HAVOC (Harmful Abstractions and Violations in Open Completions Benchmark)
measure the toxicity generated by language models across input severity and harm categories, by creating a new benchmark of open ended prefixes.
1 paper · 0 benchmarks
The HC3 (Human ChatGPT Comparison Corpus) dataset consists of nearly 40K questions and their corresponding human/ChatGPT answers.
1 paper · 0 benchmarks
The HDRT dataset is a large-scale dataset designed for infrared-guided high dynamic range (HDR) imaging.
1 paper · 0 benchmarks
HDT-QA (human driving test question answering dataset)
HDT-QA, coupled with driving manuals, offers an extensive compendium of driving instructions and driving knowledge tests across all 51 states of the US.
1 paper · 0 benchmarks
High-definition Talking Face Dataset (HDTF).
1 paper · 0 benchmarks
HEAPO (An Open Dataset for Heat Pump Optimization with Smart Electricity Meter Data and On-Site Inspection Protocols)
Heat pumps are essential for decarbonizing residential heating but consume substantial electrical energy, impacting operational costs and grid demand.
1 paper · 0 benchmarks
HEIMT (Hyperscanning EEG of Interactive Math Task)
Continuous EEG activity was recorded from each member of the dyad using an ActiveTwo head cap and the ActiveTwo Biosemi system (BioSemi, Amsterdam, Netherlands).
1 paper · 0 benchmarks
The HELMET dataset contains 910 videoclips of motorcycle traffic, recorded at 12 observation sites in Myanmar in 2016.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.