Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 175 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8353–8400 of 12,172
The Extended UCF Crime extends the UCF Crime data set that consists of 13 anomaly classes.
1 paper · 0 benchmarks
The proposed Extended-YouTube Faces (E-YTF) is an extension of the famous YouTube Faces (YTF) dataset and is specifically designed to further push the challenges of face recognition by addressing the problem of open-set face identification…
1 paper · 0 benchmarks
The dataset X of this work is an extension of the heartSeg dataset.
1 paper · 1 benchmark
214 videos under various extreme sight conditions for audiovisual repetition counting 7 vision challenges: camera viewpoint changes, cluttered background, low illumination, fast motion, disappearing activity, scale variation, low resolution
1 paper · 0 benchmarks
EyeDentify++, a dataset specifically designed for pupil diameter estimation based on webcam images and enhanced using Super Resolution techniques.
1 paper · 0 benchmarks
The EyeInfo Dataset is an open-source eye-tracking dataset created by Fabricio Batista Narcizo, a research scientist at the IT University of Copenhagen (ITU) and GN Audio A/S (Jabra), Denmark.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
FABSA (An aspect-based sentiment analysis dataset of Customer Feedback reviews)
FABSA, An aspect-based sentiment analysis dataset in the Customer Feedback space (Trustpilot, Google Play and Apple Store reviews).
1 paper · 2 benchmarks
FAD (Face Attributes Dataset)
FAD is a dataset that have roughly 200,000 attribute labels for the above traits, for over 10,000 facial images.
1 paper · 0 benchmarks
FAS100K is a large-scale visual localization dataset.
1 paper · 0 benchmarks
The FB1.5M dataset is a benchmark for Knowledge Graph Completion.
1 paper · 0 benchmarks
FB15K237-Refined is a refined version of FB15k237 by KGRefiner.
1 paper · 0 benchmarks
FBIS-22M (Field Boundary Instance Segmentation - 22M)
FBIS-22M is the largest field boundary instance segmentation dataset to date, featuring over 22 million labeled field instances across more than 672 000 high-resolution satellite image patches.
1 paper · 0 benchmarks
FCGEC (FCGEC: Fine-Grained Corpus for Chinese Grammatical Error Correction)
a fine-grained corpus to detect, identify and correct the chinese grammatical errors.
1 paper · 1 benchmark
FCoT (Foreground Chain-of-Thought)
FCoT (Chain‑of‑Thought Segmentation) is replicate the step-by-step reasoning process a human annotator follows when using SAM2 to generate masks.
1 paper · 0 benchmarks
This dataset was created for the purpose of performing NER tasks.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A large-scale isolated Indian sign language dataset.
1 paper · 1 benchmark
The data set contains point cloud data captured in an indoor environment with precise localization and ground truth mapping information.
1 paper · 0 benchmarks
FEIDEGGER (FEIDEGGER: A Multi-modal Corpus of Fashion Images and Descriptions in German)
The FEIDEGGER (fashion images and descriptions in German) dataset is a new multi-modal corpus that focuses specifically on the domain of fashion items and their visual descriptions in German.
1 paper · 0 benchmarks
Tables of the blendshapes from a group of the images of the FER2013 dataset, generated using MediaPipe library, based on the ARKit face blendshapes.
1 paper · 0 benchmarks
FETA Car-Manuals (FETA Car-Manuals dataset, image-text retrieval for foundation models' expert data performance.)
FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures.
1 paper · 2 benchmarks
FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures.
1 paper · 0 benchmarks
This dataset consists of six columns.
1 paper · 0 benchmarks
FFHQH (Flickr-Faces-HQ-Harmonization)
A new dataset for portrait harmonization based on the FFHQ.
1 paper · 1 benchmark
The FFT-75 dataset contains randomly sampled, potentially overlapping file fragments from 75 popular file types.
1 paper · 0 benchmarks
The development of the remote sensing fine-grained ship classification field necessitates large-scale realistic fine-grained ship datasets.
1 paper · 0 benchmarks
FGVD (Fine-Grained Vehicle Detection)
Fine-Grained Vehicle Detection (FGVD) is a dataset for fine-grained vehicle detection captured from a moving camera mounted on a car.
1 paper · 0 benchmarks
FGraDA (Fine-Grained Domain Adaptation Dataset)
Previous research for adapting a general neural machine translation (NMT) model into a specific domain usually neglects the diversity in translation within the same domain, which is a core problem for domain adaptation in real- world…
1 paper · 0 benchmarks
FICLE (Factual Inconsistency CLassification with Explanation)
The FICLE dataset is a derivative of the FEVER dataset, which is a collection of 185,445 claims generated by modifying sentences obtained from Wikipedia.
1 paper · 0 benchmarks
Optical images of printed circuit boards as well as detailed annotations of any text, logos, and surface-mount devices (SMDs).
1 paper · 0 benchmarks
FIG-Loneliness (FIne-Grained Loneliness) is a dataset collected by using Reddit posts in two young adult-focused forums and two loneliness related forums consisting of a diverse age group.
1 paper · 0 benchmarks
FIGRIM (FIne-GRained Image Memorability)
This is a dataset of 9428 images, 1754 of which are target images with memorability scores.
1 paper · 0 benchmarks
FIND (Fused Image dataset for convolutional neural Network-based crack Detection)
The “Fused Image dataset for convolutional neural Network-based crack Detection” (FIND) is a large-scale image dataset with pixel-level ground truth crack data for deep learning-based crack segmentation analysis.
1 paper · 0 benchmarks
FIREBALL (FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State Information)
Dungeons & Dragons (D&D) is a tabletop roleplaying game with complex natural language interactions between players and hidden state information.
1 paper · 0 benchmarks
Data used in the paper "FIRESTARTER 2: Dynamic Code Generation for Processor Stress Tests", as well as notebooks to generate plots.
1 paper · 0 benchmarks
A dataset of real-world underwater videos annotated with multi-object tracking labels.
1 paper · 0 benchmarks
FIW-MM (Families In Wild Multimedia)
A large-scale dataset for recognizing kinship in multimedia which extend FIW with multimedia data (i.e., video, audio, and contextual transcripts).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function).
1 paper · 0 benchmarks
This dataset was collected with FLOBOT - an advanced autonomous floor scrubber - includes data from four different sensors for environment perception, as well as the robot pose in the world reference frame.
1 paper · 0 benchmarks
FLOGA (wiLdfire Observations for the Greek Area)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The FLUXSynID Dataset comprises 14,889 high-resolution synthetic face identities, each uniquely represented with paired images: document-style images and trusted live-capture images.
1 paper · 0 benchmarks
FM WILN (Montmorency Forest WILN dataset)
This dataset was created while conducting the field report related to this paper.
1 paper · 0 benchmarks
FMARS (Foundation Models Annotation in Remote Sensing)
FMARS is a large-scale dataset of Very High Resolution (VHR) remote sensing images with annotations generated using Vision Foundation Models.
1 paper · 0 benchmarks
FMC-MWO2KG (The MWO2KG Failure Mode Classification Dataset)
The Failure Mode Classification dataset released in the paper "MWO2KG and Echidna: Constructing and exploring knowledge graphs from maintenance data" by Stewart et al.
1 paper · 1 benchmark
FP4S (Floor plan image segmentation via scribble-based semi-weakly-supervised learning)
We introduce a new style- and category-agnostic floor plan image parsing benchmark developed in collaboration with professional architectural designers.
1 paper · 1 benchmark
FQ-160 (Forbidden Question Dataset (160))
The forbidden question dataset they build (based on two previous works) contains 160 questions from 160 violated categories.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.