Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 70 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3313–3360 of 12,172
Thai-Chi-HD is a high resolution dataset which can be used as reference benchmark for evaluating frameworks for image animation and video generation.
9 papers · 3 benchmarks
Social media are interactive platforms that facilitate the creation or sharing of information, ideas or other forms of expression among people.
9 papers · 1 benchmark
Answering questions about why characters perform certain actions is central to understanding and reasoning about narratives.
9 papers · 0 benchmarks
UA-GEC (UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language)
UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language
9 papers · 1 benchmark
The UCF Sports dataset consists of a set of actions collected from various sports which are typically featured on broadcast television channels such as the BBC and ESPN.
9 papers · 1 benchmark
UDIVA is a new non-acted dataset of face-to-face dyadic interactions, where interlocutors perform competitive and collaborative tasks with different behavior elicitation and cognitive workload.
9 papers · 0 benchmarks
We introduce a novel Image Quality Assessment (IQA) dataset comprising 6073 UHD-1 (4K) images, annotated at a fixed width of 3840 pixels.
9 papers · 1 benchmark
UIIS (General Underwater Image Instance Segmentation dataset)
This is the first general Underwater Image Instance Segmentation (UIIS) dataset containing 4,628 images for 7 categories with pixel-level annotations for underwater instance segmentation task
9 papers · 1 benchmark
Leonardo Filipe Rodrigues Ribeiro, Pedro H.
9 papers · 1 benchmark
The Universal Morphology (UniMorph) project is a collaborative effort to improve how NLP handles complex morphology in the world’s languages.
9 papers · 1 benchmark
Have need seven multiple exposure ground truth images satisfying EV 0, ±1, ±2, ±3 for static scenes.
9 papers · 1 benchmark
VERITE (VERification of Image-TExt pairs)
Image-text claim benchmark for out-of-context detection.
9 papers · 0 benchmarks
VIL-100 is a video instance lane detection dataset, which contains 100 videos with in total 10,000 frames, acquired from different real traffic scenarios.
9 papers · 0 benchmarks
VISUELLE is a repository build upon the data of a real fast fashion company, Nunalie, and is composed of 5577 new products and about 45M sales related to fashion seasons from 2016-2019.
9 papers · 1 benchmark
VLM²-Bench: Benchmarking Vision-Language Models on Visual Cue Matching Description VLM²-Bench is the first comprehensive benchmark designed to evaluate vision-language models' (VLMs) ability to visually link matching cues across…
9 papers · 1 benchmark
VOT2020 is a Visual Object Tracking benchmark for short-term tracking in RGB.
9 papers · 1 benchmark
We present a new large-scale human value dataset called ValueNet, which contains human attitudes on 21,374 text scenarios.
9 papers · 0 benchmarks
WebCPM is a Chinese LFQA dataset.
9 papers · 0 benchmarks
WiC-TSV (Words-in-Context: Target Sense Verification)
WiC-TSV is a new multi-domain evaluation benchmark for Word Sense Disambiguation.
9 papers · 2 benchmarks
The Zenseact Open Dataset (ZOD) is a large-scale and diverse multi-modal autonomous driving (AD) dataset, created by researchers at Zenseact.
9 papers · 0 benchmarks
The iCartoonFace dataset is a large-scale dataset that can be used for two different tasks: cartoon face detection and cartoon face recognition.
9 papers · 1 benchmark
The iWildCam2020-WILDS dataset is a variant of the iWildCam 2020 dataset.
9 papers · 1 benchmark
A scholarly data set with publications’ full-text, annotated in-text citations, and links to metadata.
9 papers · 0 benchmarks
These are 10 synthetic genomics datasets generated with NEAT v3 (based on TP53 gene of Homo Sapiens) for the use case of benchmarking somatic variant callers.
8 papers · 1 benchmark
Abstract Objective This article summarizes the preparation, organization, evaluation, and results of Track 2 of the 2018 National NLP Clinical Challenges shared task.
8 papers · 0 benchmarks
Cross-source point cloud dataset for registration task.
8 papers · 0 benchmarks
3DFAW contains 23k images with 66 3D face keypoint annotations.
8 papers · 2 benchmarks
Aci-bench: a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation
8 papers · 1 benchmark
ADIMA is a novel, linguistically diverse, ethically sourced, expert annotated and well-balanced multilingual profanity detection audio dataset comprising of 11,775 audio samples in 10 Indic languages spanning 65 hours and spoken by 6,446…
8 papers · 0 benchmarks
Dataset aimed to do automated aerial scene classification of disaster events from on-board a UAV.
8 papers · 1 benchmark
AOLP (Application-oriented License Plate)
The application-oriented license plate (AOLP) benchmark database has 2049 images of Taiwan license plates.
8 papers · 2 benchmarks
APRICOT is a collection of over 1,000 annotated photographs of printed adversarial patches in public locations.
8 papers · 0 benchmarks
Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing.
8 papers · 1 benchmark
ATLAS v2.0 (Anatomical Tracings of Lesions After Stroke Dataset version 2.0)
Accurate lesion segmentation is critical in stroke rehabilitation research for the quantification of lesion burden and accurate image processing.
8 papers · 1 benchmark
ATRW (Amur Tiger Re-identification in the Wild)
The ATRW Dataset contains over 8,000 video clips from 92 Amur tigers, with bounding box, pose keypoint, and tiger identity annotations.
8 papers · 0 benchmarks
Acted Facial Expressions In The Wild (AFEW) is a dynamic temporal facial expressions data corpus consisting of close to real world environment extracted from movie
8 papers · 1 benchmark
AesBench is an expert benchmark designed to comprehensively evaluate the aesthetic perception capacities of Multimodal Large Language Models (MLLMs) when it comes to image aesthetics perception.
8 papers · 0 benchmarks
The Airport dataset is a dataset for person re-identification which consists of 39,902 images and 9,651 identities across six cameras.
8 papers · 0 benchmarks
Alibaba Cluster Trace captures detailed statistics for the co-located workloads of long-running and batch jobs over a course of 24 hours.
8 papers · 1 benchmark
AmaSum is the largest abstractive opinion summarization dataset, consisting of more than 33,000 human-written summaries for Amazon products.
8 papers · 0 benchmarks
This datasets is a subset of the Amazon reviews dataset which contain Fashion related products
8 papers · 1 benchmark
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
8 papers · 1 benchmark
Amazon-Fraud (Multi-relational Graph Dataset for Amazon Fraudulent Account Detection)
Amazon-Fraud is a multi-relational graph dataset built upon the Amazon review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
8 papers · 3 benchmarks
A large-scale, hierarchical annotated dataset of animal faces, featuring 21.9K faces from 334 diverse species and 21 animal orders across biological taxonomy.
8 papers · 0 benchmarks
The Arabic Sentiment Twitter Dataset for the Levantine dialect (ArSenTD-LEV) is a dataset of 4,000 tweets with the following annotations: the overall sentiment of the tweet, the target to which the sentiment was expressed, how the…
8 papers · 0 benchmarks
The Arena-Hard-Auto benchmark is an automatic evaluation tool for instruction-tuned Language Learning Models (LLMs)¹.
8 papers · 0 benchmarks
The Bacteria Biotope (BB) Task is part of the BioNLP Open Shared Tasks and meets the BioNLP-OST standards of quality, originality and data formats.
8 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.