Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 54 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2545–2592 of 12,172
LIVE-itw (LIVE In the Wild Image Quality Challenge Database)
Image quality assessment (IQA) databases enable researchers to evaluate the performance of IQA algorithms and contribute towards attaining the ultimate goal of objective quality assessment research - matching human perception.
14 papers · 0 benchmarks
Linux (Linux Program Dependence Graphs)
The LINUX dataset consists of 48,747 Program Dependence Graphs (PDG) generated from the Linux kernel.
14 papers · 0 benchmarks
M3KE (Massive Multi-Level Multi-Subject Knowledge Evaluation Benchmark)
M3KE is a Massive Multi-Level Multi-Subject Knowledge Evaluation benchmark, which is developed to measure knowledge acquired by Chinese large language models by testing their multitask accuracy in zero- and few-shot settings.
14 papers · 0 benchmarks
The MACHIAVELLI Benchmark is a tool designed to measure the behavior of artificial agents, particularly their ethical behavior in pursuit of their objectives¹².
14 papers · 0 benchmarks
MAP (Maybe Ambiguous Pronoun)
Maybe Ambiguous Pronoun is a dataset similar to GAP dataset, but without binary gender constraints.
14 papers · 0 benchmarks
MAVE (MAVE: : A Product Dataset for Multi-source Attribute Value Extraction)
The dataset contains 3 million attribute-value annotations across 1257 unique categories created from 2.2 million cleaned Amazon product profiles.
14 papers · 2 benchmarks
MGif is a dataset of videos containing movements of different cartoon animals.
14 papers · 1 benchmark
MHP (Multiple-Human Parsing)
The MHP dataset contains multiple persons captured in real-world scenes with pixel-level fine-grained semantic annotations in an instance-aware setting.
14 papers · 3 benchmarks
MIntRec is a novel dataset for multimodal intent recognition.
14 papers · 1 benchmark
MMPD (Multi-Domain Mobile Video Physiology Dataset)
The Multi-domain Mobile Video Physiology Dataset (MMPD), comprising 11 hours(1152K frames) of recordings from mobile phones of 33 subjects.
14 papers · 0 benchmarks
The MSRVTT-MC (Multiple Choice) dataset is a video question-answering dataset created based on the MSR-VTT dataset.
14 papers · 2 benchmarks
The dataset presents open high-resolution test clips set with different types of content: movie fragments, sport streams, live caption clips.
14 papers · 1 benchmark
MVOR (Multi-View Operating Room)
Multi-View Operating Room (MVOR) is a dataset recorded during real clinical interventions.
14 papers · 0 benchmarks
MedDG is a large-scale high-quality Medical Dialogue dataset related to 12 types of common Gastrointestinal diseases.
14 papers · 0 benchmarks
With the same format as WikiHop, the MedHop dataset is based on research paper abstracts from PubMed, and the queries are about interactions between pairs of drugs.
14 papers · 0 benchmarks
MixEval is a ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a highly capable model ranking (i.e., 0.96 correlation with Chatbot Arena) while running locally and quickly (6% the…
14 papers · 0 benchmarks
Brief Description The Neuromorphic-MNIST (N-MNIST) dataset is a spiking version of the original frame-based MNIST dataset.
14 papers · 4 benchmarks
Nutrition5k (Nutrition5k: A Comprehensive Nutrition Dataset)
Nutrition5k is a dataset of visual and nutritional data for ~5k realistic plates of food captured from Google cafeterias using a custom scanning rig.
14 papers · 0 benchmarks
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner.
14 papers · 3 benchmarks
OVBench is a benchmark tailored for real-time video understanding: - Memory, Perception, and Prediction of Temporal Contexts: Questions are framed to reference the present state of entities, requiring models to memorize/perceive/predict…
14 papers · 1 benchmark
ObjectFolder is a dataset for multisensory object-centric perception, reasoning, and interaction.
14 papers · 0 benchmarks
Open-Platypus is a family of fine-tuned and merged Large Language Models (LLMs) that achieves the strongest performance and currently stands at first place in HuggingFace's Open LLM Leaderboard.
14 papers · 0 benchmarks
OxIOD (Oxford Inertial Odometry Dataset)
OxIOD Dataset Oxford Inertial Odometry Dataset [1] is a large set of inertial data for inertial odometry which is recorded by smartphones at 100 Hz in indoor environment.
14 papers · 0 benchmarks
PANDORA is the first large-scale dataset of Reddit comments labeled with three personality models (including the well-established Big 5 model) and demographics (age, gender, and location) for more than 10k users.
14 papers · 0 benchmarks
PASS (Pictures without humAns for Self-Supervision)
PASS is a large-scale image dataset, containing 1.4 million images, that does not include any humans and which can be used for high-quality pretraining while significantly reducing privacy concerns.
14 papers · 0 benchmarks
A new large scale plane geometry problem solving dataset called PGPS9K, labeled both fine-grained diagram annotation and interpretable solution program.
14 papers · 1 benchmark
PPR10K (Portrait Photo Retouching dataset)
PPR10K is a dataset for portrait photo retouching (PPR), which aims to enhance the visual quality of a collection of flat-looking portrait photos.
14 papers · 0 benchmarks
3D video data asset of CVPR 2022 Paper "Neural 3D Video Synthesis"
14 papers · 0 benchmarks
The dataset is based on the original MNIST dataset.
14 papers · 0 benchmarks
PromptSpeech is a dataset that consists of speech and the corresponding prompts.
14 papers · 0 benchmarks
REFUGE Challenge provides a data set of 1200 fundus images with ground truth segmentations and clinical glaucoma labels, currently the largest existing one.
14 papers · 4 benchmarks
RealEstate10K is a large dataset of camera poses corresponding to 10 million frames derived from about 80,000 video clips, gathered from about 10,000 YouTube videos.
14 papers · 1 benchmark
SIBR (SIBR Dataset for VIE in the Wild)
SIBR是面向自然场景视觉信息抽取的数据集。 1)SIBR总的有1000张图片,400张测试,600张训练,包括中文、英文两种语言。 2)包含images.zip、label.zip、train.txt、test.txt四个文件,images.zip、label.zip中包含所有图片和标签,通过train.txt和test.txt区分训练和测试。…
14 papers · 1 benchmark
Dataset of clothing size variation which includes different subjects wearing casual clothing items in various sizes, totaling to approximately 2000 scans.
14 papers · 0 benchmarks
SMS-WSJ (Spatialized Multi-Speaker Wall Street Journal)
Spatialized Multi-Speaker Wall Street Journal (SMS-WSJ) consists of artificially mixed speech taken from the WSJ database, but unlike earlier databases this one considers all WSJ0+1 utterances and takes care of strictly separating the…
14 papers · 0 benchmarks
SOREL-20M is a large-scale dataset consisting of nearly 20 million files with pre-extracted features and metadata, high-quality labels derived from multiple sources, information about vendor detections of the malware samples at the time of…
14 papers · 0 benchmarks
SRRS (Snow Removal in Realistic Scenario)
SRRS (Snow Removal in Realistic Scenario) contains 15000 synthesized snow images and 1000 snow images in real scenarios downloaded from the Internet.
14 papers · 0 benchmarks
A schema-guided task-oriented dialog dataset consisting of 127,833 utterances and knowledge base queries across 5,820 task-oriented dialogs in 13 domains that is especially designed to facilitate task and domain transfer learning in…
14 papers · 0 benchmarks
SUMMIT is a high-fidelity simulator that facilitates the development and testing of crowd-driving algorithms.
14 papers · 0 benchmarks
SVTP dataset stands for Scene Text Recognition Datasets.
14 papers · 1 benchmark
ShapeGlot (ShapeGlot: Learning Language for Shape Differentiation)
ShapeGlot: Learning Language for Shape Differentiation
14 papers · 0 benchmarks
SinD (A Drone Dataset at Signalized Intersection in China)
The SIND dataset is based on 4K video captured by drones, providing information including traffic participant trajectories, traffic light status, and high-definition maps
14 papers · 0 benchmarks
Purpose Medical imaging has become increasingly important in diagnosing and treating oncological patients, particularly in radiotherapy.
14 papers · 0 benchmarks
The T2Dv2 dataset consists of 779 tables originating from the English-language subset of the WebTables corpus.
14 papers · 4 benchmarks
TAU Urban Acoustic Scenes 2019 development dataset consists of 10-seconds audio segments from 10 acoustic scenes: airport, indoor shopping mall, metro station, pedestrian street, public square, street with medium level of traffic,…
14 papers · 2 benchmarks
TLP (Track Long and Prosper)
A new long video dataset and benchmark for single object tracking.
14 papers · 0 benchmarks
This dataset was introduced by [1], but was not used in its experiment.
14 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.