Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 179 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8545–8592 of 12,172
Dataset Files The official dataset files are hosted at https://dx.doi.org/10.21227/1y7r-am78.
1 paper · 0 benchmarks
A new dataset containing over 550K pairs (covering 143 km^2 area) of RGB and aerial LIDAR depth images.
1 paper · 0 benchmarks
This is the static test data from the study "Global Geolocated Realtime Data of Interfleet Urban Transit Bus Iding" collected by GRD-TRT-BUF-4I.
1 paper · 0 benchmarks
GRND (Gramophone Recording Noise Dataset)
Dataset of noise segments extracted from gramophone recordings.
1 paper · 0 benchmarks
A scholarly named entity recognition dataset with focus on machine learning models and datasets.
1 paper · 0 benchmarks
GTA-UAV dataset provides a large continuous area dataset (covering 81.3km2) for UAV visual geo-localization, expanding the previously aligned drone-satellite pairs to arbitrary drone-satellite pairs to better align with real-world…
1 paper · 0 benchmarks
GUISS dataset (Meshes, textures, Blend files, stereo datasets, depth maps, depth estimations))
We provide all the expected data inputs to GUISS such as meshes, texture images, and blend files.
1 paper · 0 benchmarks
GUITAR-FX-DIST is a dataset of electric guitar recordings processed with overdrive, distortion, and fuzz audio effects.
1 paper · 0 benchmarks
GUITAR-FX-DIST is a dataset of electric guitar recordings processed with overdrive, distortion and fuzz audio effects.
1 paper · 0 benchmarks
GUITAR-FX-DIST is a dataset of electric guitar recordings processed with overdrive, distortion and fuzz audio effects.
1 paper · 0 benchmarks
GUITAR-FX-DIST is a dataset of electric guitar recordings processed with overdrive, distortion and fuzz audio effects.
1 paper · 0 benchmarks
Details about the creation of the dataset can be seen in https://arxiv.org/abs/2110.06139.
1 paper · 0 benchmarks
Gait3D-Parsing is a dataset for gait recognition in the wild.
1 paper · 0 benchmarks
Gambling Address Dataset is a collection of 10,423 gambling addresses that have transactions with gambling contracts.
1 paper · 0 benchmarks
Gambling Contract Dataset is a collection of 260 gambling smart contracts from decentralized gambling websites, such as Dicether, Degens.
1 paper · 0 benchmarks
GamePad that can be used to explore the application of machine learning methods to theorem proving in the Coq proof assistant.
1 paper · 0 benchmarks
GameQA is a large-scale, diverse, and challenging multimodal reasoning dataset designed to enhance the general reasoning capabilities of Vision Language Models (VLMs).
1 paper · 0 benchmarks
GameWikiSum is a domain-specific (video game) dataset for multi-document summarization, which is one hundred times larger than commonly used datasets, and in another domain than news.
1 paper · 0 benchmarks
We construct Gaze-CIFAR-10, a gaze-augmented image dataset based on the standard CIFAR-10 benchmark, enhanced with human eye-tracking annotations collected using the HTC VIVE Pro Eye headset.
1 paper · 1 benchmark
GeBiD (Geometric shapes Bimodal Dataset)
We provide a custom synthetic bimodal dataset, called GeBiD, designed specifically for the comparison of the joint- and cross-generative capabilities of Multimodal Variational Autoencoders.
1 paper · 0 benchmarks
This dataset encompasses 265 speeches (over 200,000 tokens) from the German Bundestag, primarily from the 19th legislative term (2017-2021), given by 195 distinct speakers representing 6 political parties.
1 paper · 2 benchmarks
GelSight Young's Modulus Dataset ============== by Michael Burgess Dataset of tactile images collected over grasping common objects labelled with the objects' Young's Moduli.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Learning dynamical systems that can generalize to various parameter changes in their underlying ODEs or PDEs is a significant but challenging task.
1 paper · 0 benchmarks
GenAIPABench is a specialized dataset designed to evaluate Generative AI-based Privacy Assistants (GenAIPAs).
1 paper · 0 benchmarks
dataset for Generative Photography
1 paper · 0 benchmarks
GenPlot (GenPlot: 500k pre-generated plots)
This dataset contains the pre-generated dataset referenced in the GenPlot Paper.
1 paper · 0 benchmarks
- Images for Classification, Segmentation, Object Detection, Upsampling, and Edge LLM - Feature Noise-Augmented Dataset for Semantic Communication Kindly Check: https://huggingface.co/datasets/CQILAB/GenSC-6G
1 paper · 0 benchmarks
Mapping of free text gender entries to one of three genders: Male, Female, Non-Binary.
1 paper · 0 benchmarks
https://osf.io/btjqw/?viewonly=f31eda86e7b04ac886734a26cd2ce43d
1 paper · 0 benchmarks
Genomics Adversarial Attack Sample dataset
1 paper · 0 benchmarks
GenoTEX (An LLM Agent Benchmark for Automated Gene Expression Data Analysis)
GenoTEX (Genomics Data Automatic Exploration Benchmark) is a benchmark dataset for the automated analysis of gene expression data to identify disease-associated genes while considering the influence of other biological factors.
1 paper · 0 benchmarks
The Topic-Based Paragraph Classification in Genocide-Related Court Transcripts (GTC) dataset is the first reference corpus annotated with samples from genocide tribunals in different international criminal courts.
1 paper · 0 benchmarks
Genre2Movies (Compositional queries for Movie recommendation)
Genre annotations for movies The file genre2movies.csv contains genre-movie tuples based on Wikidata annotations (https://www.wikidata.org/).
1 paper · 0 benchmarks
GeoEDdA (A Gold Standard Dataset for Geo-semantic Annotation of Diderot & d’Alembert’s Encyclopédie)
Dataset Description - Authors: Ludovic Moncla, Katherine McDonough and Denis Vigier in the framework of the GEODE project.
1 paper · 0 benchmarks
GeoJEPAD is a multimodal dataset combining OpenStreetMap (OSM) data (attributes and geometries) with high-resolution aerial imagery from diverse urban areas.
1 paper · 0 benchmarks
A simple dataset consisting of three geometric shapes (Triangle, Rectangle, Ellipsoid) of similar sizes but different orientations.
1 paper · 0 benchmarks
GeoQuestions1089 is a crowdsourced geospatial question-answering dataset that targets the Knowledge Graph YAGO2geo.
1 paper · 1 benchmark
This dataset reports counts of active GitHub contributors (activity: 2019/2020) geolocated in early 2021.
1 paper · 0 benchmarks
GerDaLIR (A German Dataset for Legal Information Retrieval)
GerDaLIR The German Dataset for Legal Information Retrieval (GerDaLIR) is a legal information retrieval dataset comprising a large collection of documents, passages and relevance labels.
1 paper · 0 benchmarks
GerMS-AT (GERMS-AT: A Sexism/Misogyny Dataset of Forum Comments from an Austrian Online Newspaper)
This dataset contains 7984 user comments from an Austrian online newspaper.
1 paper · 2 benchmarks
Dataset Description This dataset contains rental property listings scraped from Tonaton.com, one of Ghana's leading online classifieds platforms.
1 paper · 0 benchmarks
The Gigaword Entailment dataset is a dataset for entailment prediction between an article and its headline.
1 paper · 0 benchmarks
GitBugs (GitBugs: Bug Reports for Duplicate Detection, Retrieval Augmented Generation, Triage, and More)
GitBugs is a comprehensive and up-to-date dataset comprising over 150,000 bug reports from nine actively maintained open-source projects, including Firefox, Cassandra, and VS Code.
1 paper · 1 benchmark
This is a dataset of 10.6 million GitHub projects that are copies of others, and link each record with the project's ultimate parent.
1 paper · 0 benchmarks
A distant supervision dataset by linking the entire English ClueWeb09 corpus to Freebase.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.