Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 122 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5809–5856 of 12,172
The UIT-ViWikiQA is a dataset for evaluating sentence extraction-based machine reading comprehension in the Vietnamese language.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
The Sheffield (previously UMIST) Face Database consists of 564 images of 20 individuals (mixed race/gender/appearance).
3 papers · 1 benchmark
UNDD (Urban Night Driving Dataset)
UNDD consists of 7125 unlabelled day and night images; additionally, it has 75 night images with pixel-level annotations having classes equivalent to Cityscapes dataset.
3 papers · 0 benchmarks
The newly introduced UP-COUNT dataset includes drone footage captured with cameras from the DJI Mini 2 family UAV.
3 papers · 1 benchmark
UPFD-GOS (User Preference-aware Fake News Detection)
The Gossipcop variant of the UPFD dataset for benchmarking.
3 papers · 1 benchmark
The US-4 is a dataset of Ultrasound (US) images.
3 papers · 0 benchmarks
UZLF (Leuven-Haifa High-Resolution Fundus Image Dataset for Retinal Blood Vessel Segmentation and Glaucoma Diagnosis)
The Leuven-Haifa dataset contains 240 disc-centered fundus images of 224 unique patients (75 patients with normal tension glaucoma, 63 patients with high tension glaucoma, 30 patients with other eye diseases and 56 healthy controls) from…
3 papers · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
Urban is one of the most widely used hyperspectral data used in the hyperspectral unmixing study.
3 papers · 1 benchmark
UzWordnet is a lexical-semantic database, or a “word-net”, for the (Northern) Uzbek language (native: O’zbek till) compatible with Princeton Wordnet.
3 papers · 0 benchmarks
V-HICO is a dataset for human-object interaction in videos.
3 papers · 0 benchmarks
A synthetic depth estimation dataset for benchmark rendered from a high-quality CAD indoor environment - About 3.5K RGBD pairs with left-right stereo - Challenging viewing direction - Challenging different light condition
3 papers · 1 benchmark
VATEX Adverbs is a subset from VATEX with extracted verb-adverb annotations.
3 papers · 2 benchmarks
VBR (VBR: A Vision Benchmark in Rome)
This dataset presents a vision and perception research dataset collected in Rome, featuring RGB data, 3D point clouds, IMU, and GPS data.
3 papers · 0 benchmarks
VID Dataset (The Visual-Inertial-Dynamical Multirotor Dataset)
The Visual-Inertial-Dynamical (VID) dataset not only focuses on traditional six degrees of freedom (6-DOF) pose estimation, but also provides dynamical characteristics of the flight platform for external force perception or dynamics-aided…
3 papers · 0 benchmarks
Vehicle-Rear is a novel dataset for vehicle identification that contains more than three hours of high-resolution videos, with accurate information about the make, model, color and year of nearly 3,000 vehicles, in addition to the position…
3 papers · 0 benchmarks
a vessel dataset using 85 videos.
3 papers · 1 benchmark
ViMQ is a Vietnamese dataset of medical questions from patients with sentence-level and entity-level annotations for the Intent Classification and Named Entity Recognition tasks.
3 papers · 0 benchmarks
This dataset is used for spam review detection (opinion spam reviews) on Vietnamese E-commerce website
3 papers · 0 benchmarks
Video Localized Narratives is a new form of multimodal video annotations connecting vision and language.
3 papers · 0 benchmarks
VideoMatting108 is a large-scale video matting and trimap generation dataset with 80 training and 28 validation foreground video clips with ground-truth alpha mattes.
3 papers · 0 benchmarks
Hugging Face Datasets (New!) | Website | Github Repository | arXiv e-Print The Visual Writing Prompts (VWP) dataset contains almost 2K selected sequences of movie shots, each including 5-10 images.
3 papers · 0 benchmarks
Source: paper Visual Question Answering (VQA) is the task of returning the answer to a question about an image.
3 papers · 0 benchmarks
The Vocal Folds dataset is a dataset for automatic segmentation of laryngeal endoscopic images.
3 papers · 0 benchmarks
Voice conversion (VC) is a technique to transform a speaker identity included in a source speech waveform into a different one while preserving linguistic information of the source speech waveform.
3 papers · 0 benchmarks
VoxSim is a perceptual voice similarity dataset created to develop perceptual speaker similarity evaluation systems.
3 papers · 0 benchmarks
WDC-Dialogue is a dataset built from the Chinese social media to train EVA.
3 papers · 0 benchmarks
WFDD (Woven Fabric Defect Detection)
WFDD is a dataset for benchmarking anomaly detection methods with a focus on textile inspection.
3 papers · 1 benchmark
WHU-Specular is a large dataset of annotated specular highlight regions created from real-world images.
3 papers · 0 benchmarks
In our benchmark WHYSHIFT, we explore distribution shifts on 5 real-world tabular datasets from the economic and traffic sectors with natural spatiotemporal distribution shifts.We only pick 7 typical settings out of 22 settings and select…
3 papers · 0 benchmarks
WLD (WildLife Documentary)
WildLife Documentary is an animal object detection dataset.
3 papers · 0 benchmarks
This shared task will examine automatic evaluation metrics for machine translation.
3 papers · 0 benchmarks
Lev Finkelstein, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, and Eytan Ruppin.
3 papers · 1 benchmark
Test-driven benchmark to challenge LLMs to write JavaScript React application GitHub Script
3 papers · 1 benchmark
WebUAV-3M is a new million-scale Unmanned Aerial Vehicle (UAV) tracking benchmark consisting of 4,485 videos with more than 3M frames from the Internet.
3 papers · 0 benchmarks
The WebVid-CoVR dataset is a collection of video-text-video triplets that can be used for the task of composed video retrieval (CoVR).
3 papers · 1 benchmark
Dataset Description The dataset described in the provided text is focused on social media polls collected from Weibo, a popular Chinese microblogging platform.
3 papers · 3 benchmarks
WiLI-2018 is a benchmark dataset for monolingual written natural language identification.
3 papers · 0 benchmarks
Wiki-Convert is a 900,000+ sentences dataset of precise number annotations from English Wikipedia.
3 papers · 0 benchmarks
A dataset of single-sentence edits crawled from Wikipedia.
3 papers · 0 benchmarks
The WikiScenes dataset consists of paired images and language descriptions capturing world landmarks and cultural sites, with associated 3D models and camera poses.
3 papers · 0 benchmarks
WikiText-TL-39 is a benchmark language modeling dataset in Filipino that has 39 million tokens in the training set.
3 papers · 0 benchmarks
WikiWiki is a dataset for understanding entities and their place in a taxonomy of knowledge—their types.
3 papers · 0 benchmarks
Wikipedia Title is a dataset for learning character-level compositionality from the character visual characteristics.
3 papers · 0 benchmarks
Wireless AI Research Dataset is a flexible and easy-to-use dataset with realistic environments designed for various wireless AI tasks.
3 papers · 0 benchmarks
This dataset refers to the two images acquired by the WorldView-2 satellite, representing Miami.
3 papers · 1 benchmark
A multilingual dataset for the task of multilingual claim span identification.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.