Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 96 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4561–4608 of 12,172
The corpus represents the largest existing corpus of Catalan containing 687 million words, which is a significant increase given that until now the biggest corpus of Catalan, CuCWeb, counts 166 million words.
5 papers · 0 benchmarks
iPer is a new dataset, with diverse styles of clothes in videos, for the evaluation of human motion imitation, appearance transfer, and novel view synthesis.
5 papers · 0 benchmarks
Created from endoscopic video feeds of real-world surgical procedures.
5 papers · 0 benchmarks
mTVR is a large-scale multilingual video moment retrieval dataset, containing 218K English and Chinese queries from 21.8K TV show video clips.
5 papers · 0 benchmarks
7,672 human written natural language navigation instructions for routes in OpenStreetMap with a focus on visual landmarks.
5 papers · 2 benchmarks
The dataset consists in many runs of the same quantum circuit on different IBM quantum machines.
5 papers · 0 benchmarks
The tStoryCloze refers to the "Topic StoryCloze" benchmark, which is a spoken version of the StoryCloze textual benchmark.
5 papers · 0 benchmarks
The TweetSentBR Dataset is a valuable resource for sentiment analysis in Brazilian Portuguese.
5 papers · 1 benchmark
voraus-AD contains machine data of a collaborative robot, which moves a can by performing an industrial pick-and-place task.
5 papers · 1 benchmark
word2word contains easy-to-use word translations for 3,564 language pairs.
5 papers · 0 benchmarks
The $O2$Perm dataset is created from the Membrane Society of Australasia portal.
4 papers · 0 benchmarks
Description: 1,995 People Face Images Data (Asian race).
4 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
Data set of 360-degree equirectangular videos, gaze recordings, eye movement (EM) ground-truth and an automatic EM classification algorithm.
4 papers · 0 benchmarks
Provides a large-scale synthetic dataset which contains accurate ground truth depth of various photo-realistic scenes.
4 papers · 0 benchmarks
We established a 3D evaluation benchmark, 3D MM-Vet, to assess the 4-level capacity in embodied interaction scenarios, varying from basic perception to control statements generation.
4 papers · 1 benchmark
To collect the 3D Vehicle Tracking Simulation Dataset, a driving simulation is used to obtain accurate 3D bounding box annotations at no cost of human efforts.
4 papers · 0 benchmarks
The 3DNet dataset is a free resource for object class recognition and 6DOF pose estimation from point cloud data.
4 papers · 0 benchmarks
The 8TAGS dataset is a corpus specifically created for the evaluation of sentence representations in Polish.
4 papers · 0 benchmarks
A Game Of Sorts is a collaborative image ranking task.
4 papers · 0 benchmarks
ACM (Association for Computing Machinery
Active Contour Model
algebraic collective model
and-Compare Module
Active Contour Models)
The ACM dataset contains papers published in KDD, SIGMOD, SIGCOMM, MobiCOMM, and VLDB and are divided into three classes (Database, Wireless Communication, Data Mining).
4 papers · 1 benchmark
ADE-OoD is a public benchmark for dense out-of-distribution detection in general natural images.
4 papers · 1 benchmark
This dataset is described in the ALTA 2021 Shared Task website and associated CodaLab competition.
4 papers · 0 benchmarks
ALTO (Aerial-view Large-scale Terrain-Oriented)
ALTO is a vision-focused dataset for the development and benchmarking of Visual Place Recognition and Localization methods for Unmanned Aerial Vehicles.
4 papers · 0 benchmarks
AMA (Articulated Mesh Animation)
Articulated Mesh Animation (AMA) is a real-world dataset containing 10 mesh sequences depicting 3 different humans performing various actions
4 papers · 0 benchmarks
ARAUS (Affective Responses to Augmented Urban Soundscapes)
Choosing optimal maskers for existing soundscapes to effect a desired perceptual change via soundscape augmentation is non-trivial due to extensive varieties of maskers and a dearth of benchmark datasets with which to compare and develop…
4 papers · 0 benchmarks
ARC-DA (ARC Direct Answer Questions)
ARC Direct Answer Questions (ARC-DA) dataset consists of 2,985 grade-school level, direct-answer ("open response", "free form") science questions derived from the ARC multiple-choice question set released as part of the AI2 Reasoning…
4 papers · 0 benchmarks
ARD-16 (Ati Real-world Dataset)
We create ARD-16 (Ati Realworld Dataset), a first of its kind real-world paired correspondence dataset, by applying our dataset generation method on 16-beam VLP-16 Puck LiDAR scans on a slow-moving Unmanned Ground Vehicle.
4 papers · 0 benchmarks
ARMBench is a large-scale, object-centric benchmark dataset for robotic manipulation in the context of a warehouse.
4 papers · 1 benchmark
ATIS (vi) (Vietnamese Intent Detection and Slot Filling)
This is a dataset for intent detection and slot filling for the Vietnamese language.
4 papers · 2 benchmarks
Advising Corpus is a dataset based on an entirely new collection of dialogues in which university students are being advised which classes to take.
4 papers · 1 benchmark
AeroRIT is a hyperspectral dataset to facilitate aerial hyperspectral scene understanding.
4 papers · 0 benchmarks
The DeepMind Alchemy environment is a meta-reinforcement learning benchmark that presents tasks sampled from a task distribution with deep underlying structure.
4 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
Amazon Fine Foods is a dataset that consists of reviews of fine foods from amazon.
4 papers · 0 benchmarks
AndroidHowTo contains 32,436 data points from 9,893 unique How-To instructions and split into training (8K), validation (1K) and test (900).
4 papers · 0 benchmarks
AnimeRun is a 2D animation visual correspondence dataset.
4 papers · 0 benchmarks
In this paper, we present AnlamVer, which is a semantic model evaluation dataset for Turkish designed to evaluate word similarity and word relatedness tasks while discriminating those two relations from each other.
4 papers · 0 benchmarks
ArCOV19-Rumors is an Arabic COVID-19 Twitter dataset for misinformation detection composed of tweets containing claims from 27th January till the end of April 2020.
4 papers · 0 benchmarks
Data set covering a set of debatable topics, where for each topic and stance, a set of triplets of the form is provided.
4 papers · 0 benchmarks
The Argoverse 2 Lidar Dataset is a collection of 20,000 scenarios with lidar sensor data, HD maps, and ego-vehicle pose.
4 papers · 0 benchmarks
AutoChart is a dataset for chart-to-text generation, a task that consists on generating analytical descriptions of visual plots.
4 papers · 0 benchmarks
A whole-body FDG-PET/CT dataset with manually annotated tumor lesions (FDG-PET-CT-Lesions) 1,014 studies (900 patients)
4 papers · 0 benchmarks
AwA Pose is a large scale animal keypoint dataset with ground truth annotations for keypoint detection of quadruped animals from images.
4 papers · 0 benchmarks
Breast cancer is the most common invasive cancer in women, affecting more than 10% of women worldwide.
4 papers · 0 benchmarks
BCCD is a small-scale dataset for blood cells detection.
4 papers · 0 benchmarks
As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, we present the BIOSCAN-5M Insect dataset to the machine learning community.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.