Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 57 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2689–2736 of 12,172
SciEval is a comprehensive and multi-disciplinary evaluation benchmark designed to assess the performance of large language models (LLMs) in the scientific domain.
13 papers · 0 benchmarks
A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
13 papers · 0 benchmarks
Story Commonsense is a new large-scale dataset with rich low-level annotations and establishes baseline performance on several new tasks, suggesting avenues for future research.
13 papers · 0 benchmarks
Super-CLEVR is a dataset for Visual Question Answering (VQA) where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently.
13 papers · 0 benchmarks
SynWoodScape (Synthetic Surround-view Fisheye Camera Dataset for Autonomous Driving)
SynWoodScape is a synthetic version of the surround-view dataset covering many of its weaknesses and extending it.
13 papers · 0 benchmarks
TCR (Temporal and Causal Reasoning dataset)
A dataset of Joint Reasoning for Temporal and Causal Relations
13 papers · 0 benchmarks
In order to create the TED-talks dataset, 3,035 YouTube videos were downloaded using the "TED talks" query.
13 papers · 1 benchmark
TTPLA (Transmission Towers and Power Lines (TTPLA))
TTPLA is a public dataset which is a collection of aerial images on Transmission Towers (TTs) and Power Lines (PLs).
13 papers · 0 benchmarks
The TUT Acoustic Scenes 2017 dataset is a collection of recordings from various acoustic scenes all from distinct locations.
13 papers · 1 benchmark
The COLOSSEUM (The COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation)
To realize effective large-scale, real-world robotic applications, we must evaluate how well our robot policies adapt to changes in environmental conditions.
13 papers · 1 benchmark
This dataset encompasses a diverse range of tactile features that are instrumental in bifurcating various material properties.
13 papers · 0 benchmarks
UFDD (Unconstrained Face Detection Dataset)
Unconstrained Face Detection Dataset (UFDD) aims to fuel further research in unconstrained face detection.
13 papers · 0 benchmarks
UPFD (User Preference-aware Fake News Detection)
For benchmarking, please refer to its variant UPFD-POL and UPFD-GOS.
13 papers · 0 benchmarks
Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them.
13 papers · 1 benchmark
The VQA-CP dataset was constructed by reorganizing VQA v2 such that the correlation between the question type and correct answer differs in the training and test splits.
13 papers · 1 benchmark
ViQuAE is a dataset for KVQAE (Knowledge-based Visual Question Answering about named Entities), a task which consists in answering questions about named entities grounded in a visual context using a Knowledge Base.
13 papers · 0 benchmarks
Vidore (Visual Document Retrieval Benchmark)
It is collection regrouping all datasets constituting the ViDoRe benchmark.
13 papers · 0 benchmarks
Visual Madlibs is a dataset consisting of 360,001 focused natural language descriptions for 10,738 images.
13 papers · 0 benchmarks
WSVD (Web Stereo Video Dataset)
The Web Stereo Video Dataset consists of 553 stereoscopic videos from YouTube.
13 papers · 0 benchmarks
The Watch-n-Patch dataset was created with the focus on modeling human activities, comprising multiple actions in a completely unsupervised setting.
13 papers · 0 benchmarks
A Multi-Task 4D Radar-Camera Fusion Dataset for Autonomous Driving on Water Surfaces description of the dataset WaterScenes, the first multi-task 4D radar-camera fusion dataset on water surfaces, which offers data from multiple sensors,…
13 papers · 2 benchmarks
A benchmark dataset for data-driven medium-range weather forecasting, a topic of high scientific interest for atmospheric and computer scientists alike.
13 papers · 0 benchmarks
Who-did-What collects its corpus from news and provides options for questions similar to CBT.
13 papers · 0 benchmarks
YUP++ (YUP++ Dynamic Scenes dataset)
A new and challenging video database of dynamic scenes that more than doubles the size of those previously available.
13 papers · 1 benchmark
Yelp-Fraud (Multi-relational Graph Dataset for Yelp Spam Review Detection)
Yelp-Fraud is a multi-relational graph dataset built upon the Yelp spam review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
13 papers · 3 benchmarks
This is a dataset for evaluating summarisation methods for research papers.
13 papers · 3 benchmarks
This dataset contains 98k 2-hop explanations for questions in the QASC dataset, with annotations indicating if they are valid (~25k) or invalid (~73k) explanations.
13 papers · 0 benchmarks
Ecoset, an ecologically motivated image dataset, is a large-scale image dataset designed for human visual neuroscience, which consists of over 1.5 million images from 565 basic-level categories.
13 papers · 0 benchmarks
openai.com/blog/safety-gym/
13 papers · 0 benchmarks
A popular dataset for node classification on heterogeneous graphs.
12 papers · 1 benchmark
The ACNE04 dataset includes 3756 Chinese face images with Acne.
12 papers · 1 benchmark
ADAM (Adam: automatic detection challenge on age-related macular degeneration)
ADAM is organized as a half day Challenge, a Satellite Event of the ISBI 2020 conference in Iowa City, Iowa, USA.
12 papers · 1 benchmark
The dataset contains product information from AliExpress Sports & Entertainment category.
12 papers · 2 benchmarks
Aesthetic Visual Analysis is a dataset for aesthetic image assessment that contains over 250,000 images along with a rich variety of meta-data including a large number of aesthetic scores for each image, semantic labels for over 60…
12 papers · 1 benchmark
AndroZoo is a growing collection of Android apps collected from several sources, including the official Google Play app market and a growing collection of various metadata of those collected apps aiming at facilitating the Android-relevant…
12 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
12 papers · 1 benchmark
BenchLMM (BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models)
Large Multimodal Models (LMMs) such as GPT-4V and LLaVA have shown remarkable capabilities in visual reasoning with common image styles.
12 papers · 1 benchmark
Bongard-HOI testifies to which extent your few-shot visual learner can quickly induce the true HOI concept from a handful of images and perform reasoning with it.
12 papers · 1 benchmark
The BotNet dataset is a set of topological botnet detection datasets forgraph neural networks.
12 papers · 0 benchmarks
This paper introduces the Broad Twitter Corpus (BTC), which is not only significantly bigger, but sampled across different regions, temporal periods, and types of Twitter users.
12 papers · 2 benchmarks
CDCP (Cornell eRulemaking Corpus)
The Cornell eRulemaking Corpus – CDCP is an argument mining corpus annotated with argumentative structure information capturing the evaluability of arguments.
12 papers · 3 benchmarks
CDR (BioCreative V CDR Task Corpus)
The BioCreative V CDR task corpus is manually annotated for chemicals, diseases and chemical-induced disease (CID) relations.
12 papers · 2 benchmarks
CICERO (Contextualized Commonsense Inference in Dialogues)
CICERO contains 53,000 inferences for five commonsense dimensions -- cause, subsequent event, prerequisite, motivation, and emotional reaction -- collected from 5600 dialogues.
12 papers · 4 benchmarks
CIRCLE is a dataset containing 10 hours of full-body reaching motion from 5 subjects across nine scenes, paired with ego-centric information of the environment represented in various forms, such as RGBD videos.
12 papers · 1 benchmark
The COCO-MLT is created from MS COCO-2017, containing 1,909 images from 80 classes.
12 papers · 2 benchmarks
COMPAS (machine bias risk assessments in criminal sentencing)
Dataset used by ProPublica to assess and analyse the fairness of the COMPAS (Correctional Offender Management Profiling for Alternative Sanctions) software.
12 papers · 0 benchmarks
Along with COVID-19 pandemic we are also fighting an infodemic'.
12 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.