Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 225 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10753–10800 of 12,172
A high-resolution version of VGGFace2 for academic face editing purposes.
1 paper · 0 benchmarks
A dataset of gas dispersion simulations in complex 3D environments.
1 paper · 0 benchmarks
Gaze following attracts much attention recently while existing databases commonly lack audio information.
1 paper · 0 benchmarks
The VIA dataset is a dataset for aiding the visually impaired.
1 paper · 0 benchmarks
VILT (Video Instructions Linking for Complex Tasks)
VILT is a new benchmark collection of tasks and multimodal video content.
1 paper · 0 benchmarks
VIRDO Dataset (VIRDO Simulated Kitchen Utensil Deformation Dataset)
From https://github.com/MMintLab/VIRDO/blob/master/data/datasetreadme.txt, 1.
1 paper · 0 benchmarks
A visible-light and thermal-infrared images dataset for dual-spectrum depth estimation.
1 paper · 0 benchmarks
VISEM-Tracking is a dataset consisting of 20 video recordings of 30s of spermatozoa with manually annotated bounding-box coordinates and a set of sperm characteristics analyzed by experts in the domain.
1 paper · 0 benchmarks
VISO (VIdeo Satellite Objects)
This dataset is a large-scale dataset for moving object detection and tracking in satellite videos, which consists of 40 satellite videos captured by Jilin-1 satellite platforms.
1 paper · 0 benchmarks
VISOR is a dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A novel dataset that consists of content-based and video-specific features extracted from publicly available scientific video lectures and several metrics related to user engagement.
1 paper · 0 benchmarks
VLUC (Video-Like Urban Computing)
VLUC (Video-Like Urban Computing) is a benchmark for video-like computing on citywide traffic density and crowd prediction.
1 paper · 0 benchmarks
VMD (Virtual Moderation Dataset)
This dataset contains synthetically generated discussions and annotations using exclusively Large Language Model (LLM) agents.
1 paper · 0 benchmarks
VME & CDSI (Vehicles in the Middle East (VME) & Car Detection in Satellite Imagery (CDSI) datasets)
Vehicles in the Middle East (VME) dataset, designed explicitly for vehicle detection in high-resolution satellite images from Middle Eastern countries.
1 paper · 1 benchmark
VNEMOS (Vietnamese Speech Emotion Dataset)
This research introduces the dataset that we created to test voice emotional recognition models with Vietnamese data.
1 paper · 0 benchmarks
VOCEdits: A benchmark for precise geometric object-level editing Sample format: (input image, edit prompt, input mask, ground-truth output mask, ...)
1 paper · 0 benchmarks
VOT2015 (Visual Object Tracking Challenge 2015)
VOT2015 is a visual object tracking dataset.
1 paper · 0 benchmarks
The benchmark for VPData, the largest video inpainting dataset, which comprises over 390K clips (> 866.7 hours) and features precise masks and detailed video captions.
1 paper · 0 benchmarks
The largest video inpainting dataset comprises over 390K clips (> 866.7 hours), featuring precise masks and detailed video captions.
1 paper · 0 benchmarks
VQA-MHUG is a 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker.
1 paper · 0 benchmarks
VQA-OV (Visual Quality Assessment of Omnidirectional Video)
Collects 60 reference sequences and 540 impaired sequences.
1 paper · 0 benchmarks
The datasets includes curves drawn on 3D surfaces (triangle meshes) in Virtual Reality.
1 paper · 0 benchmarks
Data used for the paper Combining Motion Matching and Orientation Prediction to Animate Avatars for Consumer-Grade VR Devices.
1 paper · 0 benchmarks
VR-Folding contains garment meshes of 4 categories from CLOTH3D dataset, namely Shirt, Pants, Top and Skirt.
1 paper · 0 benchmarks
This dataset collection includes three files used for the experiments.
1 paper · 0 benchmarks
VSLID (Very Small Lego Image Dataset)
VSLID stands for Very Small Lego Image Dataset.
1 paper · 0 benchmarks
VSTaR-1M is a 1M instruction tuning dataset, created using Video-STaR, with the source datasets: Kinetics700 STAR-benchmark FineDiving The videos for VSTaR-1M can be found in the links above.
1 paper · 0 benchmarks
VTQA (Visual Text Question Answering)
VTQA is a dataset containing open-ended questions about image-text pairs.
1 paper · 0 benchmarks
Combines CoVaxFrames and HpVaxFrames into a unified dataset of 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines and 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines.
1 paper · 0 benchmarks
A Natural Language Resource for Learning to Recognize Misinformation about the COVID-19 and HPV Vaccines.
1 paper · 0 benchmarks
Validity and Novelty are determined in a comparative setting between two conclusions at a time.
1 paper · 1 benchmark
ValiMath is a high-quality benchmark consisting of 2,147 carefully curated mathematical questions designed to evaluate an LLM's ability to verify the correctness of math questions based on multiple logic-based and structural criteria.
1 paper · 0 benchmarks
This dataset contains a total of 11 variables.
1 paper · 0 benchmarks
AlgorithmComparison: Comparison of algorithms on benchmark test cases.
1 paper · 0 benchmarks
Presented stimuli of the listening experiment performed in the course of the master's thesis: K.
1 paper · 0 benchmarks
Various URL Datasets These are collections of URLs for benchmarking purposes.
1 paper · 0 benchmarks
The Vashantor dataset consists of 32,500 sentences from different regions, including Chittagong, Noakhali, Sylhet, Barishal, and Mymensingh.
1 paper · 0 benchmarks
Vastextures (Vast Dataset for textures and PBR materials)
VasTexture is a free giant repository of textures and PBR materials extracted from real-world images.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
VedantaNY-10M is a curated dataset of over 750 hours of transcripts from public discourses on the Indian philosophy of Advaita Vedanta.
1 paper · 0 benchmarks
We present a new collection of 1,981 Vega-Lite specifications, which is used to demonstrate the generalizability and viability of our NL generation framework.
1 paper · 0 benchmarks
VerbCL is a dataset that consists of the citation graph of court opinions, which cite previously published court opinions in support of their arguments.
1 paper · 0 benchmarks
Verifee is a dataset of news articles with fine-grained trustworthiness annotations.
1 paper · 0 benchmarks
Verified Smart Contracts Code Comments is a dataset of real Ethereum smart contract functions, containing "code, comment" pairs of both Solidity and Vyper source code.
1 paper · 1 benchmark
Verified Smart Contracts is a dataset of real Ethereum smart contracts, containing both Solidity and Vyper source code.
1 paper · 0 benchmarks
Verireason-RTL-Coder7breasoningtb For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log…
1 paper · 0 benchmarks
Verireason-RTL-Coder7breasoningtbsimple For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.