Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 132 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6289–6336 of 12,172
FINDSum (Financial Report Document Summarization)
FINDSum is a large-scale dataset for long text and multi-table summarization.
2 papers · 0 benchmarks
FKD (Football Keywords Dataset)
The football keyword dataset (FKD), as a new keyword spotting dataset in Persian, is collected with crowdsourcing.
2 papers · 1 benchmark
FPDS (Fallen People Data Set)
A benchmark for detecting fallen people lying on the floor.
2 papers · 0 benchmarks
FRGC-Morphs is a dataset of morphed faces selected from the publicly available FRGC dataset [1].
2 papers · 0 benchmarks
A large-scale and accurate dataset for vision-based railway traffic light detection and recognition.The recordings were made on selected running trains in France and benefited from carefully hand-labeled annotations.
2 papers · 0 benchmarks
FSVOD-500 is a large-scale video dataset comprising of 500 classes with class-balanced videos in each category for few-shot learning.
2 papers · 0 benchmarks
FSVQA (Full-Sentence Visual Question Answering)
Full-Sentence Visual Question Answering (FSVQA) dataset, consisting of nearly 1 million pairs of questions and full-sentence answers for images, built by applying a number of rule-based natural language processing techniques to original…
2 papers · 0 benchmarks
FaMoS (Facial Motion across Subjects)
FaMoS is a dynamic 3D head dataset from 95 subjects, each performing 28 motion sequences.
2 papers · 0 benchmarks
FaceOcc is a high-quality face occlusion dataset which contains all mislabeled occlusions in CelebAMask-HQ and complements some occlusions and textures from the internet.
2 papers · 0 benchmarks
This failure dataset contains information on the events collected in the OpenStack cloud computing platform during three different campaigns of fault-injection experiments performed with three different workloads.
2 papers · 0 benchmarks
FairPrism is a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harms relating to gender and sexuality.
2 papers · 0 benchmarks
FarsBase-KBP contains 22015 sentences, in which the entities and relation types are linked to the FarsBase ontology.
2 papers · 0 benchmarks
Fashion-MNT is large-scale bilingual product description dataset called Fashion-MMT, which contains over 114k noisy and 40k manually cleaned description translations with multiple product images.
2 papers · 0 benchmarks
We provide multiple human annotations for each test image in Fashion-MNIST.
2 papers · 0 benchmarks
📄 Read 💾 Code 🔗 Webpage 💻 Demo 🤗 Huggingface Dataset 💬 Discussions Overview Users interact with QA systems and leave feedback.
2 papers · 0 benchmarks
Fetoscopic Placental Vessel Segmentation and Registration (FetReg2021) challenge was organized as part of the MICCAI2021 Endoscopic Vision (EndoVis) challenge.
2 papers · 0 benchmarks
The fetoscopy placenta dataset is associated with our MICCAI2020 publication titled “Deep Placental Vessel Segmentation for Fetoscopic Mosaicking”.
2 papers · 0 benchmarks
FinBench is a benchmark for evaluating the performance of machine learning models with both tabular data inputs and profile text inputs.
2 papers · 0 benchmarks
This dataset enriches the benchmark Room-to-Room (R2R) dataset by dividing the instructions into sub-instructions and pairing each of those with their corresponding viewpoints in the path.
2 papers · 0 benchmarks
This dataset includes all music sources, background noises and impulse-reponses (IR) samples and conversation speech that have been used in the work "Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning"…
2 papers · 0 benchmarks
FinnSentiment introduces a 27,000 sentence dataset (in Finnish) annotated independently with sentiment polarity by three native annotators.
2 papers · 0 benchmarks
Schools of inland silversides (Menidia beryllina, n=14 individuals per school) were recorded in the Lauder Lab at Harvard University while swimming at 15 speeds (0.5 to 8 BL/s, body length, at 0.5 BL/s intervals) in a flow tank with a…
2 papers · 1 benchmark
FixaTons is a large collection of datasets human scanpaths (temporally ordered sequences of fixations) and saliency maps.
2 papers · 1 benchmark
The Flickr 8k Audio Caption Corpus contains 40,000 spoken captions of 8,000 natural images.
2 papers · 0 benchmarks
Given a sentence in the source language, generate a translation in the target language.
2 papers · 0 benchmarks
FormulaNet FormulaNet is a new large-scale Mathematical Formula Detection dataset.
2 papers · 0 benchmarks
This dataset is made up of forward-looking sonar images containing ten classes of underwater debris.
2 papers · 1 benchmark
FracAtlas (A Dataset for Fracture Classification, Localization and Segmentation of Musculoskeletal Radiographs)
FractureAtlas is a musculoskeletal bone fracture dataset with annotations for deep learning tasks like classification, localization, and segmentation.
2 papers · 0 benchmarks
This dataset is dialog dataset collected in a Wizard-of-Oz fashion.
2 papers · 0 benchmarks
An object-centric dataset consiting of 52 RGB sequences of cars
2 papers · 0 benchmarks
The Freiburg Spatial Relations dataset features 546 scenes each containing two out of 25 household objects.
2 papers · 0 benchmarks
FuLG is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully extracted from Common Crawl.
2 papers · 0 benchmarks
Fusion-DHL is a multimodal sensor dataset with ground-truth positions.
2 papers · 0 benchmarks
G-VUE (General-purpose Visual Understanding Evaluation)
General-purpose Visual Understanding Evaluation (G-VUE) is a comprehensive benchmark covering the full spectrum of visual cognitive abilities with four functional domains -- Perceive, Ground, Reason, and Act.
2 papers · 0 benchmarks
GAFA (Gaze from afar dataset)
We introduce a new dataset of annotated surveillance videos of freely moving people taken from a distance in both indoor and outdoor scenes.
2 papers · 0 benchmarks
GATITOS (Google's Additional Translations Into Tail-languages: Often Short)
The GATITOS (Google's Additional Translations Into Tail-languages: Often Short) dataset is a high-quality, multi-way parallel dataset of tokens and short phrases, intended for training and improving machine translation models.
2 papers · 0 benchmarks
GBUSV (Gallbladder Ultrasound Videos)
Description GBUSV is a un-annotated dataset consisting of ultrasound videos of of patients with either of a malignant or a non-malignant gallbladder.
2 papers · 0 benchmarks
GEMRec-18K is a dense prompt-model interaction dataset that consists of 18K images generated by pairing 200 generative models with 90 prompts collected from real-world usages.
2 papers · 0 benchmarks
GEOBench-VLM, a comprehensive benchmark specifically designed to evaluate VLMs on geospatial tasks, including scene understanding, object counting, localization, fine-grained categorization, and temporal analysis.
2 papers · 0 benchmarks
GIRT-Data (GitHub Issue Report Template Dataset)
GIRT-Data is the first and largest dataset of issue report templates (IRTs) in both YAML and Markdown format.
2 papers · 0 benchmarks
GIS (Github Issue Similarity)
This dataset can be used for semantic textual similarity tasks.
2 papers · 0 benchmarks
Grammatical error correction dataset for text from Yahoo!
2 papers · 0 benchmarks
Golos is a Russian speech dataset suitable for speech research.
2 papers · 0 benchmarks
We release the dataset for non-commercial research.
2 papers · 0 benchmarks
GPT-generated and hum-written academic abstract corpus with over 600k samples in Computer Science, Physics, and Humanity Science.
2 papers · 0 benchmarks
GPI corpus (Government Privacy Instructions Corpus)
The GPI Corpus is a collection of 1,043 privacy laws, regulations, and guidelines ("GPIs") covering 182 jurisdictions around the world.
2 papers · 0 benchmarks
GQA-OOD is a new dataset and benchmark for the evaluation of VQA models in OOD (out of distribution) settings.
2 papers · 0 benchmarks
This is the latest version of our datasets, and is built upon GTA-V for expressive human pose and shape estimation.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.