Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 148 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7057–7104 of 12,172
Underwater Trash Detection Dataset Overview The Underwater Trash Detection Dataset is a custom-annotated dataset designed to address the challenges of underwater trash detection caused by varying environmental features.
2 papers · 0 benchmarks
The Unified SSL Benchmark (USB) consists of 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio) to evaluate self-supervised learning (SSL) methods.
2 papers · 0 benchmarks
UofTPed50 is an object detection and tracking dataset which uses GPS to ground truth the position and velocity of a pedestrian.
2 papers · 0 benchmarks
This corpus was constructed by collecting 10,008 reviews from various domains, including sports, food, software, politics, and entertainment.
2 papers · 1 benchmark
V2VBench is a comprehensive benchmark designed to evaluate video editing methods.
2 papers · 0 benchmarks
To test interpolation performance on various texture types, we developed a new test set, VFITex, which contains twenty 100-frame UHD or HD videos at 24, 30 or 50 FPS, collected from the Xiph, Mitch Martinez Free 4K Stock Footage, UVG…
2 papers · 1 benchmark
Vision-based Fallen Person (VFP290K) is a novel, large-scale dataset for the detection of fallen persons composed of fallen person images collected in various real-world scenarios.
2 papers · 1 benchmark
325 word images intended for font recognition, whose fonts are included in [VFR-447] (and [VFR-2420]).
2 papers · 1 benchmark
VGG-Sound Sync is an audio-visual synchronisation benchmark based on videos collected from YouTube.
2 papers · 0 benchmarks
VGaokao is a verification style reading comprehension dataset designed for native speakers' evaluation.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
The dataset, VIST-Edit, includes 14,905 human-edited versions of 2,981 machine-generated visual stories.
2 papers · 0 benchmarks
VNDS (VNDS: A Vietnamese Dataset for Summarization)
A single-document Vietnamese summarization dataset
2 papers · 1 benchmark
VOICe is a dataset for the development and evaluation of domain adaptation methods for sound event detection.
2 papers · 0 benchmarks
Language Identification Dataset
2 papers · 2 benchmarks
VQA 360° is a dataset for visual question answering on 360° images containing around 17,000 real-world image-question-answer triplets for a variety of question types.
2 papers · 0 benchmarks
VQDv1 (Visual Query Detection v1)
In Visual Query Detection (VQD), a system is given a query (prompt) natural language and an image, and then the system must produce 0 - N boxes that satisfy that query.
2 papers · 0 benchmarks
The dataset contains traffic traces collected from 3 different VR applications.
2 papers · 0 benchmarks
Data used for the paper SparsePoser: Real-time Full-body Motion Reconstruction from Sparse Data It contains over 1GB of high-quality motion capture data recorded with an Xsens Awinda system while using a variety of VR applications in Meta…
2 papers · 0 benchmarks
VTC (Videos, Titles and Comments)
VTC is a large-scale multimodal dataset containing video-caption pairs (~300k) alongside comments that can be used for multimodal representation learning.
2 papers · 0 benchmarks
A dataset for Visual Voice Activity Detection extracted from the LRS3 dataset.
2 papers · 0 benchmarks
Binary labels for Validity and Novelty respectively are given for each Conclusion.
2 papers · 1 benchmark
ValueConsistency is a dataset of both controversial and uncontroversial questions in English, Chinese, German, and Japanese for topics from the U.S., China, Germany, and Japan.
2 papers · 0 benchmarks
The code to create the dataset is available here.
2 papers · 2 benchmarks
VerilogEval Dataset The VerilogEval Dataset is a benchmark specifically designed to assess the ability of large language models (LLMs) to generate syntactically correct and functionally accurate Verilog code.
2 papers · 1 benchmark
ViText2SQL is a dataset for the Vietnamese Text-to-SQL semantic parsing task, consisting of about 10K question and SQL query pairs.
2 papers · 0 benchmarks
Vibrating Plates (Vibrating Plates Dataset for Vibroacoustic Frequency Response Prediction)
We present a structured benchmark dataset for a representative vibroacoustic problem: Predicting the frequency response for vibrating plates.
2 papers · 0 benchmarks
VideoForensicsHQ is a benchmark dataset for face video forgery detection, providing high quality visual manipulations.
2 papers · 0 benchmarks
Large-scale benchmark dataset of full-field digital mammography, called VinDr-Mammo, which consists of 5,000 four-view exams with breast-level assessment and finding annotations.
2 papers · 0 benchmarks
The Virtual Gallery dataset is a synthetic dataset that targets multiple challenges such as varying lighting conditions and different occlusion levels for various tasks such as depth estimation, instance segmentation and visual…
2 papers · 0 benchmarks
The Vistas-NP dataset is an out-of-distribution detection dataset based on the Mapillary Vistas dataset.
2 papers · 0 benchmarks
A large-scale multi-view RGBD visual affordance learning dataset, a benchmark of 47210 RGBD images from 37 object categories, annotated with 15 visual affordance categories and 35 cluttered/complex scenes with different objects and…
2 papers · 0 benchmarks
Visual Beliefs is a dataset of abstract scenes to study visual beliefs.
2 papers · 0 benchmarks
The dataset of the paper: Dataset and Case Studies for Visual Near-Duplicates Detection in the Context of Social Media'', by Hana Matatov, Mor Naaman, and Ofra Amir.
2 papers · 0 benchmarks
A dataset containing 5000 images with 37,993 thousand relationships.
2 papers · 0 benchmarks
Dataset for visual servoing (VS) and camera pose estimation.
2 papers · 0 benchmarks
VizWiz-Priv includes 8,862 regions showing private content across 5,537 images taken by blind people.
2 papers · 0 benchmarks
A large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering.
2 papers · 0 benchmarks
VocBench is a framework that benchmark the performance of state-of-the art neural vocoders.
2 papers · 0 benchmarks
WDC SOTAB is a benchmark that features two annotation tasks: Column Type Annotation and Columns Property Annotation.
2 papers · 2 benchmarks
WDC-PAVE (Web Data Commones - Product Attribute Value Extraction)
The datasets contains 1,420 human annotated product offers, systematically selected from the Web Data Commons Product Matching Corpus, featuring 24,582 annotated attribute-value pairs, making it a valuable resource for both product…
2 papers · 1 benchmark
COVID19 Data from the World Health Organization
2 papers · 1 benchmark
WT-WT (Will-They-Won't-They)
Will-They-Won't-They (WT-WT) is a large dataset of English tweets targeted at stance detection for the rumor verification task.
2 papers · 0 benchmarks
A benchmark for Human-Human Interaction (HHI) recognition as free text.
2 papers · 0 benchmarks
The Wallhack1.8k dataset comprises 1,806 CSI amplitude spectrograms (and raw WiFi packet time series) corresponding to three activity classes: "no presence," "walking," and "walking + arm-waving." WiFi packets were transmitted at a…
2 papers · 0 benchmarks
Ward2ICU is a vital signs dataset of inpatients from the general ward.
2 papers · 0 benchmarks
Wastewater catchment area data are essential for wastewater treatment capacity planning and have recently become critical for operationalising wastewater-based epidemiology (WBE) for COVID-19.
2 papers · 0 benchmarks
One of the founding fathers of marine mammal bioacoustics, William Watkins, carried out pioneering work with William Schevill at the Woods Hole Oceanographic Institution for more than four decades, laying the groundwork for our field today.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.