Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 224 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10705–10752 of 12,172
We present UniTalk, a novel dataset specifically designed for the task of active speaker detection, emphasizing challenging scenarios to enhance model generalization.
1 paper · 0 benchmarks
Unidecor (A unified deception corpus for cross-corpus deception detection.)
UNIDECOR is a unified corpus consolidating publicly available textual deception datasets into a common format.
1 paper · 0 benchmarks
The dataset is hosted on GitHub: https://github.com/ucsdsysnet/nsl-empirical-analysis and explained in our paper.
1 paper · 0 benchmarks
UniRef90 is generated by clustering UniRef100 seed sequences.
1 paper · 0 benchmarks
Uniswap (Replication Data for: Uniswap Daily Transaction Indices by Network)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Unitail (The United Retail Datasets)
The United Retail Datasets (Unitail) is a large-scale benchmark of basic visual tasks on products that challenges algorithms for detecting, reading, and matching.
1 paper · 0 benchmarks
The United-Syn-Med dataset is a specialized medical speech dataset designed to evaluate and improve Automatic Speech Recognition (ASR) systems within the healthcare domain.
1 paper · 0 benchmarks
A package for creating Unity Perception compatible synthetic people.
1 paper · 0 benchmarks
The University of Washington/Northwestern University (UW/NU) Corpus contains recordings and textgrids of Pacific Northwest and Northern Cities speakers reading a subset of the IEEE "Harvard" sentences.
1 paper · 0 benchmarks
Unpaired dataset: The dataset is built by ourselves, and there are all real haze images from websites.
1 paper · 0 benchmarks
Rolling Shutter dataset captured with Unreal engine
1 paper · 0 benchmarks
The Unsplash Dataset is made up of over 350,000+ contributing global photographers and data sourced from hundreds of millions of searches across a nearly unlimited number of uses and contexts.
1 paper · 0 benchmarks
Unsplash2K is high-resolution image dataset with 2K resolution.
1 paper · 0 benchmarks
Inpainting networks are typically benchmarked on samples from Places2 dataset.
1 paper · 0 benchmarks
Unsynchronized dynamic blender dataset for multi-view dynamic NeRFs for evaluating MAE between the predicted time offsets and the ground truth.
1 paper · 0 benchmarks
The prospective upper body thermal images SARS-CoV2 association study was designed to test the hypothesis that thermal videos may aid in the early diagnosis of COVID-19.
1 paper · 0 benchmarks
Urban Dict spelling variant is a variant spelling dataset for use of NLP research in the informal domain.
1 paper · 0 benchmarks
A main goal of the Urban Soundscapes of the World project is to create a reference database of examples of urban acoustic environments, consisting of high-quality immersive audiovisual recordings (360-degree video and spatial audio), in…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is the translation of the MS-marco dataset, marking it the first large-scale urdu IR dataset.
1 paper · 0 benchmarks
Urdu News Headlines Dataset with VOA and BBC An Urdu news headlines dataset is a collection of news headlines in the Urdu language, typically scraped from news websites and social media platforms.
1 paper · 1 benchmark
The UrduDoc Dataset is a benchmark dataset for Urdu text line detection in scanned documents.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset includes User Story (or Issue) text descriptions, User Story titles, and Story Points from 33 software development projects, comprising a total of 20,479 User Stories (or issues) extracted from GitLab repositories, amounting…
1 paper · 0 benchmarks
Despite the successes of recent developments in visual AI, different shortcomings still exist; from missing exact logical reasoning, to abstract generalization abilities, to understanding complex and noisy scenes.
1 paper · 0 benchmarks
V-MIND enhanced the MIND dataset with news pictures.
1 paper · 0 benchmarks
V-PCCD (simulated Visual Point Cloud Change Detection dataset)
A simulated dataset built in Unreal Engine 4 with AirSim.
1 paper · 0 benchmarks
V-Rank (SIMULATED AIRCRAFT TRAJECTORY FOR THEORETICAL VELOCITY RANKING)
Abstract This data set is a data set used for aircraft theoretical velocity ranking.
1 paper · 0 benchmarks
V2AIX (A Multi-Modal Real-World Dataset of ETSI ITS V2X Messages in Public Road Traffic)
Connectivity is a main driver for the ongoing megatrend of automated mobility: future Cooperative Intelligent Transport Systems (C-ITS) will connect road vehicles, traffic signals, roadside infrastructure, and even vulnerable road users,…
1 paper · 0 benchmarks
V3C1 (the Vimeo Creative Commons Collection 1)
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, and will serve as evaluation basis for the Video Browser Showdown 2019-2021 and TREC Video Retrieval…
1 paper · 0 benchmarks
VADD (Video Anomaly Detection Dataset (VADD))
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
VAST Absorption is a dataset of spatial binaural features annotated with acoustic properties such as the 3D source position and the walls’ absorption coefficients.
1 paper · 0 benchmarks
Uses same clean speech as VoiceBank+Demand but more noise types.
1 paper · 1 benchmark
VCAS-Motion (Video Class Agnostic Segmentation Benchmark)
Video class agnostic segmentation (VCAS) is the task of segmenting objects without regards to its semantics combining appearance, motion and geometry from monocular video sequences.
1 paper · 0 benchmarks
VCG+112K (Video Instruction Dataset 112K)
Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs.
1 paper · 0 benchmarks
This task stems from the observation that text embedded in images is intrinsically different from common visual elements and natural language due to the need to align the modalities of vision, text, and text embedded in images.
1 paper · 0 benchmarks
VD-Ref is a dataset with ground-truth mappings from both noun phrases and pronouns to image regions.
1 paper · 0 benchmarks
VDQG (Visual Discriminative Question Generation)
The Visual Discriminative Question Generation (VDQG) dataset contains 11202 ambiguous image pairs collected from Visual Genome.
1 paper · 0 benchmarks
VESSEL12 (VESsel SEgmentation in the Lung 2012)
1 paper · 0 benchmarks
VESUS (Varied Emotion in Syntactically Uniform Speech)
The Varied Emotion in Syntactically Uniform Speech (VESUS) repository is a lexically controlled database collected by the NSA lab.
1 paper · 0 benchmarks
VETRA is a dataset for vehicle tracking in aerial image sequences and presents unique challenges such as low frame rates, small and fast-moving objects, as well as high camera movement.
1 paper · 0 benchmarks
VFD-2000 is a video fight detection dataset containing more than 2000 videos.
1 paper · 0 benchmarks
A synthetic dataset containing word images of 447 typefaces with font variations for each typeface, created for visual font recognition.
1 paper · 1 benchmark
A synthetic dataset containing 447 typefaces with only one font variation for each typeface, created for visual font recognition.
1 paper · 1 benchmark
Enable visual relation detection and serves as an extension to Visual Genome (VG).
1 paper · 0 benchmarks
The VGG Cell dataset (made up entirely of synthetic images) is the main public benchmark used to compare cell counting techniques.
1 paper · 0 benchmarks
The VGG Physical Property dataset introduced in the paper is a set of dataset containing positive/negative pairs for understanding of different physical properties, including scene geometry, material, support relation, shadow, occlusion…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.