Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 80 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3793–3840 of 12,172
The People's Speech is a free-to-download 30,000-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA (with a CC-BY subset).
7 papers · 0 benchmarks
Enriches the TimeML annotations of TimeBank by adding information about the Topic Time in terms of Klein (1994).
7 papers · 2 benchmarks
TransCG is the first large-scale real-world dataset for transparent object depth completion and grasping, which contains 57,715 RGB-D images of 51 transparent objects and many opaque objects captured from different perspectives (~240…
7 papers · 1 benchmark
UBI-Fights - Concerning a specific anomaly detection and still providing a wide diversity in fighting scenarios, the UBI-Fights dataset is a unique new large-scale dataset of 80 hours of video fully annotated at the frame level.
7 papers · 2 benchmarks
UHRSD (Ultra High-Resolution Saliency Detection Dataset)
Recent salient object detection (SOD) methods based on deep neural network have achieved remarkable performance.
7 papers · 1 benchmark
UIT-ViCTSD (UIT Vietnamese Constructive and Toxic Speech Detection)
UIT-ViCTSD (Vietnamese Constructive and Toxic Speech Detection) is a dataset for constructive and toxic speech detection in Vietnamese.
7 papers · 0 benchmarks
UPAR (Unified Pedestrian Attribute Recognition)
The Task: The challenge will use an extension of the UPAR Dataset [1], which consists of images of pedestrians annotated for 40 binary attributes.
7 papers · 1 benchmark
UPIQ (Unified Photometric Image Quality)
Contains over 4,000 images created by realigning and merging existing HDR and standard-dynamic-range (SDR) datasets.
7 papers · 0 benchmarks
This dataset contains 70 (30 falls + 40 activities of daily living) sequences.
7 papers · 0 benchmarks
A chemical synthesis route dataset constructed from the USPTO reaction dataset (1976-Sep2016) and a list of commercially available building blocks from eMolecules (~23.1M molecules).
7 papers · 1 benchmark
This dataset was collected with the goal of assessing dialog evaluation metrics.
7 papers · 1 benchmark
This dataset was collected with the goal of assessing dialog evaluation metrics.
7 papers · 1 benchmark
VIVOS is a free Vietnamese speech corpus consisting of 15 hours of recording speech prepared for Automatic Speech Recognition task.
7 papers · 1 benchmark
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
7 papers · 1 benchmark
VPCD (Video Person-Clustering)
VPCD contains multi-modal annotations (face, body and voice) for all primary and secondary characters from a range of diverse TV-shows and movies.
7 papers · 1 benchmark
This dataset provides a new split of VQA v2 (similarly to VQA-CP v2), which is built of questions that are hard to answer for biased models.
7 papers · 1 benchmark
ViHOS (Hate Speech Spans Detection for Vietnamese)
The first human-annotated corpus containing 26k spans on 11k comments
7 papers · 0 benchmarks
VidHOI is a video-based human-object interaction detection benchmark.
7 papers · 2 benchmarks
VoicePrivacy 2020 is a dataset for developing anonymization solutions for speech technology.
7 papers · 0 benchmarks
WALT (Watch and Learn TimeLapse Images)
We introduce a new dataset, Watch and Learn Time-lapse (WALT), consisting of multiple (4K and 1080p) cameras capturing urban environments over a year.
7 papers · 1 benchmark
SOTAB V2 features two annotation tasks: Column Type Annotation (CTA) and Columns Property Annotation (CPA).
7 papers · 2 benchmarks
The WeChat dataset for fake news detection contains more than 20k news labelled as fake news or not.
7 papers · 1 benchmark
A multivariate spatio-temporal benchmark dataset for meteorological forecasting based on real-time observation data from ground weather stations.
7 papers · 16 benchmarks
WebFG-496 is a dataset for fine-grained recognition that contains 200 subcategories of the "Bird" (Web-bird), 100 subcategories of the Aircraft" (Web-aircraft), and 196 subcategories of the "Car" (Web-car).
7 papers · 0 benchmarks
WikiCatSum is a domain specific Multi-Document Summarisation (MDS) dataset.
7 papers · 0 benchmarks
An annotated dataset of 1m crowd-sourced annotations that cover 100k talk page diffs (with 10 judgements per diff) for personal attacks, aggression, and toxicity.
7 papers · 0 benchmarks
The WikiTables-TURL dataset was constructed by the authors of TURL and is based on the WikiTable corpus, which is a large collection of Wikipedia tables.
7 papers · 3 benchmarks
XSafety is the first multilingual safety benchmark specifically designed for Large Language Models (LLMs).
7 papers · 0 benchmarks
YCB-Slide (YCB-Slide: A tactile interaction dataset)
The YCB-Slide dataset comprises of DIGIT sliding interactions on YCB objects.
7 papers · 0 benchmarks
YT-BB (YouTube-BoundingBoxes)
YouTube-BoundingBoxes (YT-BB) is a large-scale data set of video URLs with densely-sampled object bounding box annotations.
7 papers · 1 benchmark
YUD+ (Additional Vanishing Point Labels for the York Urban Database)
YUD+ is a dataset containing additional Vanishing Point Labels for the York Urban Database.
7 papers · 0 benchmarks
This dataset is a new knowledge-base (KB) of hasPart relationships, extracted from a large corpus of generic statements.
7 papers · 0 benchmarks
iPhone dataset is a challenging benchmarks for dynamic reconstruction.
7 papers · 1 benchmark
k-qa (K-QA: A Real-World Medical Q&A Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
7 papers · 0 benchmarks
The m2cai16-tool-locations dataset contains spatial tool annotations for 2,532 frames across the first 10 videos in the m2cai16-tool dataset, which includes 15 videos in total.
7 papers · 0 benchmarks
A Modular Simulation Framework and Benchmark for Robot Learning.
7 papers · 0 benchmarks
2D HeLa is a dataset of fluorescence microscopy images of HeLa cells stained with various organelle-specific fluorescent dyes.
6 papers · 0 benchmarks
300-VW (300 Videos in the Wild)
300 Videos in the Wild (300-VW) is a dataset for evaluating facial landmark tracking algorithms in the wild.
6 papers · 2 benchmarks
This is a synthetic dataset constructed to stimulate the development and evaluation of 3D lane detection methods.
6 papers · 0 benchmarks
A large-scale synthetic dataset with 2.5 Million photo-realistic images of 80 subjects performing 70 activities and wearing diverse outfits.
6 papers · 0 benchmarks
The 3DSeg-8 is a collection of several publicly available 3D segmentation datasets from different medical imaging modalities, e.g.
6 papers · 0 benchmarks
ADVANCE (AuDio Visual Aerial sceNe reCognition datasEt)
The AuDio Visual Aerial sceNe reCognition datasEt (ADVANCE) is a brand-new multimodal learning dataset, which aims to explore the contribution of both audio and conventional visual messages to scene recognition.
6 papers · 0 benchmarks
AH36M (Ambiguous Human3.6M)
Since H36M is captured in a controlled environment, it rarely depicts challenging real-world scenarios such as body occlusions that are the main source of ambiguity in the single-view 3D shape estimation problem.
6 papers · 1 benchmark
AM-2K (Animal Matting 2,000 Dataset)
AM-2k (Animal Matting 2,000 Dataset) consists of 2,000 high-resolution images collected and carefully selected from websites with open licenses.
6 papers · 1 benchmark
The AMI Meeting Corpus is a multi-modal data set comprising 100 hours of meeting recordings.
6 papers · 1 benchmark
AO-CLEVr is a new synthetic-images dataset containing images of "easy" Attribute-Object categories, based on the CLEVr.
6 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.