Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 92 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4369–4416 of 12,172
This is a dataset for a shot boundary detection task.
5 papers · 1 benchmark
MUSIC-AVQA v2.0 balances the original MUSIC-AVQA dataset in each QA category and sub-category.
5 papers · 1 benchmark
MarKG (Multimodal analogical reasoning Knowledge Graph)
The MarKG dataset has 11,292 entities, 192 relations and 76,424 images, including 2,063 analogy entities and 27 analogy relations.
5 papers · 0 benchmarks
The Market1501-Attributes dataset is built from the Market1501 dataset.
5 papers · 1 benchmark
a multidimensional AVR benchmark with 770 puzzles composed of six core knowledge patterns, geometric and abstract shapes, and five different task configurations.
5 papers · 0 benchmarks
MatSynth MatSynth is a Physically Based Rendering (PBR) materials dataset designed for modern AI applications.
5 papers · 0 benchmarks
Understanding what makes a video memorable has a very broad range of current applications, e.g., education and learning, content retrieval and search, content summarization, storytelling, targeted advertising, content recommendation and…
5 papers · 0 benchmarks
MediaSpeech is a media speech dataset (you might have guessed this) built with the purpose of testing Automated Speech Recognition (ASR) systems performance.
5 papers · 1 benchmark
A large, realistic multimodal dataset consisting of real personal photos and crowd-sourced questions/answers.
5 papers · 1 benchmark
The Middlebury 2006 is a stereo dataset of indoor scenes with multiple handcrafted layouts.
5 papers · 0 benchmarks
MindCraft is a fine-grained dataset of collaborative tasks performed by pairs of human subjects in the 3D virtual blocks world of Minecraft.
5 papers · 0 benchmarks
This is the dataset used by the automatic sparse attention compression method MoA.
5 papers · 0 benchmarks
MolOpt (Molecular Optimization)
Open-source benchmark for Practical Molecular Optimization (PMO), to facilitate the transparent and reproducible evaluation of algorithmic advances in molecular optimization.
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
MulRan (MulRan: Multimodal Range Dataset for Urban Place Recognition)
MulRan is a dataset for Place Recognition and SLAM.
5 papers · 0 benchmarks
Multi-IF (Multi-turn and multilingual instruction following)
We introduce Multi-IF, a new benchmark designed to assess LLMs' proficiency in following multi-turn and multilingual instructions.
5 papers · 0 benchmarks
Introudced from Multimodal Meme Dataset (MultiOFF) for Identifying Offensive Content in Image and Text
5 papers · 1 benchmark
N-Digit MNIST is a multi-digit MNIST-like dataset.
5 papers · 0 benchmarks
NELA-GT-2019 is an updated version of the NELA-GT-2018 dataset.
5 papers · 0 benchmarks
NERDS 360 (NeRF for Reconstruction, Decomposition and Scene Synthesis of 360° outdoor scenes)
We present a large-scale dataset for 3D urban scene understanding.
5 papers · 0 benchmarks
NFCorpus is a full-text English retrieval data set for Medical Information Retrieval.
5 papers · 1 benchmark
NISP (NITK-IISc Multilingual Multi-accent Speaker Profiling)
This dataset contains speech recordings along with speaker physical parameters (height, weight, shoulder size, age ) as well as regional information and linguistic information.
5 papers · 0 benchmarks
NLPeer is a multidomain corpus of more than 5k papers and 11k review reports from five different venues.
5 papers · 0 benchmarks
The NTIRE 2021 HDR was built for the first challenge on high-dynamic range (HDR) imaging that was part of the New Trends in Image Restoration and Enhancement (NTIRE) workshop, held in conjunction with CVPR 2021.
5 papers · 0 benchmarks
The NYU Symmetry database contains 176 single-symmetry and 63 multiple-symmetry images (.png files) with accompanying ground-truth annotations (.mat files).
5 papers · 0 benchmarks
NYU-VP is a new dataset for multi-model fitting, vanishing point (VP) estimation in this case.
5 papers · 0 benchmarks
NaturalCodeBench (NCB) is a comprehensive code benchmark designed to mirror the complexity and variety of scenarios in real coding tasks¹².
5 papers · 0 benchmarks
This is a catalogue and repository of network datasets with the aim of aiding scientific research.
5 papers · 0 benchmarks
Synthetically Generated Night-time Weather Degraded Database
5 papers · 1 benchmark
NucMM is a dataset for segmenting 3D cell nuclei from microscopy image volumes that pushes the task forward to the sub-cubic millimeter scale.
5 papers · 0 benchmarks
Contains 13.6k masked-word-prediction probes, 10.5k for fine-tuning and 3.1k for testing.
5 papers · 0 benchmarks
A popular dataset for node classification on heterogeneous graphs.
5 papers · 1 benchmark
OAK (Objects Around Krishna)
OAK is a dataset for online continual object detection benchmark with an egocentric video dataset.
5 papers · 0 benchmarks
The OLR 2021 dataset contains the data for the Oriental Language Recognition (OLR) 2021 Challenge, which intends to improve the performance of language recognition systems and speech recognition systems within multilingual scenarios.
5 papers · 0 benchmarks
The ObjectsRoom dataset is based on the MuJoCo environment used by the Generative Query Network [4] and is a multi-object extension of the 3d-shapes dataset.
5 papers · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
SMAC+ offensive complicated scenario with sequential episodic buffer
5 papers · 1 benchmark
SMAC+ offensive distant scenario with sequential episodic buffer
5 papers · 1 benchmark
SMAC+ offensive hard scenario with sequential episodic buffer
5 papers · 1 benchmark
SMAC+ offensive near scenario with sequential episodic buffer
5 papers · 1 benchmark
SMAC+ offensive superhard scenario with sequential episodic buffer
5 papers · 1 benchmark
OlympicArena is a benchmark to evaluate advanced capabilities of language models across a broad spectrum of Olympic-level challenges.
5 papers · 0 benchmarks
This dataset was collected with research funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie Grant Agreement No 691025.
5 papers · 0 benchmarks
OneStopQA provides an alternative test set for reading comprehension which alleviates these shortcomings and has a substantially higher human ceiling performance.
5 papers · 0 benchmarks
OntoEvent is a new ED dataset with event correlations.
5 papers · 0 benchmarks
OoDIS (Anomaly Instance Segmentation Benchmark)
OoDIS is a benchmark dataset for anomaly instance segmentation, crucial for autonomous vehicle safety.
5 papers · 2 benchmarks
Given two entities, generating a coherent sentence describing the relation between them.
5 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.