Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 243 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11617–11664 of 12,172
This dataset contains human-annotated sense identifiers for 2562 contexts of 20 words used in the RUSSE'2018 shared task on Word Sense Induction and Disambiguation for the Russian language.
0 papers · 0 benchmarks
H3D (Humans in 3D) is a dataset of annotated people.
0 papers · 0 benchmarks
Towards automated analysis of large environments, hyperspectral sensors must be adapted into a format where they can be operated from mobile robots.
0 papers · 0 benchmarks
H²O is an image dataset annotated for Human-to-human-or-object interaction detection.
0 papers · 0 benchmarks
The International Cardiac Arrest REsearch consortium (I-CARE) Database includes baseline clinical information and continuous electroencephalogram (EEG) and electrocardiogram (ECG) recordings from comatose patients following cardiac arrest.
0 papers · 0 benchmarks
The IC-BIN dataset was introduced by Doumanoglou et al.
0 papers · 0 benchmarks
The IC-MI dataset, introduced by Tejani et al., is part of the Benchmark for 6D Object Pose Estimation (BOP).
0 papers · 0 benchmarks
ICConv (A Large-scale Automated Intent-oriented and Context-aware Conversational Search Dataset)
The dataset contains 105,811 information-seeking conversations converted from MS MARCO.
0 papers · 0 benchmarks
ICT-3DHP is collected using the Microsoft Kinect sensor and contains RGB images and depth maps of about 14k frames, divided in 10 sequences.
0 papers · 0 benchmarks
Can you detect fraud from customer transactions?
0 papers · 0 benchmarks
Overview The IITKGPFence dataset is designed for tasks related to fence-like occlusion detection, defocus blur, depth mapping, and object segmentation.
0 papers · 0 benchmarks
IKEA 3D is a dataset of IKEA 3D models and aligned images, which is suitable for pose estimation.
0 papers · 0 benchmarks
IMO (Independently Moving Objects)
Dataset of annotated independently moving objects (IMO).
0 papers · 0 benchmarks
IMO-AG-30 refers to a set of 30 classical geometry problems adapted from the International Mathematical Olympiad (IMO) contests.
0 papers · 0 benchmarks
A large paroxysmal atrial fibrillation long-term electrocardiogram monitoring database Abstract Atrial fibrillation (AF) is the most common sustained heart arrhythmia in adults.
0 papers · 0 benchmarks
This is a Dataset for Arabic/English text detection and optical character recognition.
0 papers · 0 benchmarks
The ISIBengaliCharacter dataset contains 158 classes of Bengali numerals, characters or their parts.
0 papers · 0 benchmarks
ISMIR2004 is an audio dataset consisting of 6 genres with 729 excerpts of 30 seconds.
0 papers · 0 benchmarks
This is a social interaction dataset between two subjects.
0 papers · 0 benchmarks
Im-Promptu Visual Analogy Suite is a meta-learning framework.
0 papers · 0 benchmarks
A database of images with measured probabilities that each picture will be remembered after a single view.
0 papers · 0 benchmarks
InLUT3D (Indoor Lodz University of Technology Point Cloud Dataset)
This dataset called Indoor Lodz University of Technology Point Cloud Dataset (InLUT3D) is a point cloud set tailored for real object classification and both semantic and instance segmentation tasks.
0 papers · 0 benchmarks
Perception systems of autonomous vehicles are susceptible to occlusion, especially when examined from a vehicle-centric perspective.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 5000+ original India food images captured and crowdsourced from over 800+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals at DC…
0 papers · 0 benchmarks
This dataset is collected by DataCluster Labs.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 20,000+ original Number plate images captured and crowdsourced from over 700+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals…
0 papers · 0 benchmarks
Introduction The dataset consists of Indian traffic signs images for classification and detection.
0 papers · 0 benchmarks
This dataset is collected by Datacluster Labs.
0 papers · 0 benchmarks
IndirectRequests is an LLM-generated dataset of user utterances in a task-oriented dialogue setting where the user does not directly specify their preferred slot value.
0 papers · 0 benchmarks
This dataset contains 4,403 Indonesian tweets that have been labeled into five emotion classes: love, anger, sadness, joy, and fear.
0 papers · 0 benchmarks
InfiniteRep is a synthetic, open-source dataset for fitness and physical therapy (PT) applications.
0 papers · 0 benchmarks
Infinity AI's Spills Basic Dataset is a synthetic, open-source dataset for safety applications.
0 papers · 0 benchmarks
InpaintCOCO is a benchmark to understand fine-grained concepts in multimodal models (vision-language) similar to Winoground.
0 papers · 0 benchmarks
Inria building dataset contains 360 images (5120×5120) collected from 5 cities (Austin, Chicago, Kitsap, Tyrol, and Vienna)
0 papers · 0 benchmarks
The Interestingness dataset contains movie excerpts and key-frames and corresponding ground truth files based on classification into interesting and non-interesting samples.
0 papers · 0 benchmarks
The Remote Sensing dataset contains the following key features for each annotated marking: Marking Type: Specifies whether the marking is a lane-use arrow or a crosswalk.
0 papers · 0 benchmarks
JTES (Japanese Twitter-based Emotional Speech)
We designed an emotional speech database that can be used for emotion recognition as well as recognition and synthesis of speech with various emotions.
0 papers · 0 benchmarks
JigsawPlan contains room layouts and floorplans for 98,780 single-story houses/apartments from a production pipeline, designed for the Extreme Structure from Motion (E-SfM) problem.
0 papers · 0 benchmarks
We introduce the KAIST multi-spectral dataset, which covers a greater range of drivable regions, from urban to residential, for autonomous systems.
0 papers · 0 benchmarks
The aim of KFall dataset is to contribute technology development for elderly fall detection and injury prevention.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.