Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 134 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6385–6432 of 12,172
This is a dataset consisting of complete network traces comprising benign and malicious traffic, which is feature-rich and publicly available.
2 papers · 0 benchmarks
The ICL-NUIM dataset aims at benchmarking RGB-D, Visual Odometry and SLAM algorithms.
2 papers · 0 benchmarks
The full IFCNet dataset currently consists of 19,000 CAD models distributed over 65 classes according to the taxonomy of the Industry Foundation Classes (IFC) standard.
2 papers · 1 benchmark
IG-3.5B-17k is an internal Facebook AI Research dataset for training image classification models.
2 papers · 0 benchmarks
IHDS (Indian Human Developement Survey)
IHDS is a nationally representative, multi-topic panel survey of 41,554 households in 1503 villages and 971 urban neighborhoods across India.
2 papers · 0 benchmarks
IISc VEED-Dynamic (Indian Institute of Science Virtual Environment Exploration Database for Dynamic Scenes)
IISc VEED-Dynamic consists of 200 diverse indoor and outdoor scenes (see samples below).
2 papers · 0 benchmarks
IKEA Object State Dataset is a new dataset that contains IKEA furniture 3D models, RGBD video of the assembly process, the 6DoF pose of furniture parts and their bounding box.
2 papers · 0 benchmarks
Imitation learning field requires expert data to train agents in a task.
2 papers · 0 benchmarks
IMEMNET (Image-MusicEmotion-Matching-Net)
The Image-MusicEmotion-Matching-Net (IMEMNet) dataset is a dataset for continuous emotion-based image and music matching.
2 papers · 0 benchmarks
Bearing acceleration data from three run-to-failure experiments on a loaded shaft.
2 papers · 0 benchmarks
The INRIA-Horse dataset consists of 170 horse images and 170 images without horses.
2 papers · 0 benchmarks
A significant challenge in removing shadows from indoor scenes is obtaining shadow-free images.
2 papers · 2 benchmarks
INSTRE is a benchmark for INSTance-level visual object REtrieval and REcognition (INSTRE).
2 papers · 1 benchmark
The IRFL dataset consists of idioms, similes, and metaphors with matching figurative and literal images, as well as two novel tasks of multimodal figurative understanding and preference.
2 papers · 2 benchmarks
The ISIC 2017 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
2 papers · 0 benchmarks
Sign languages are the primary means of communication for a large number of people worldwide.
2 papers · 0 benchmarks
The ISOT Fake News dataset is a compilation of several thousands fake news and truthful articles, obtained from different legitimate news sites and sites flagged as unreliable by Politifact.com.
2 papers · 0 benchmarks
ITB (Informative Tracking Benchmark)
Informative Tracking Benchmark (ITB) is a small and informative tracking benchmark with 7% out of 1.2 M frames of existing and newly collected datasets, which enables efficient evaluation while ensuring effectiveness.
2 papers · 1 benchmark
Iconary dataset is for testing multimodal communication with drawings and text.
2 papers · 0 benchmarks
Icons-50 is a dataset for studying surface variation robustness.
2 papers · 0 benchmarks
Im4Sketch is a large-scale dataset with shape-oriented set of classes for image-to-sketch generalization .
2 papers · 1 benchmark
ImDrug is a comprehensive benchmark with an open-source Python library which consists of 4 imbalance settings, 11 AI-ready datasets, 54 learning tasks and 16 baseline algorithms tailored for imbalanced learning.
2 papers · 0 benchmarks
Replication Data for: Integrating Earth Observation Data into Causal Inference: Challenges and Opportunities Details: YandWmat.csv contains individual-level observational data.
2 papers · 0 benchmarks
This ImageNet-100 dataset was introduced in the following paper, Vaze, S., Han, K., Vedaldi, A.
2 papers · 0 benchmarks
ImageNet3D for general-purpose object-level 3D understanding.
2 papers · 0 benchmarks
A Point Cloud Dataset for place recognition provided by PointNetVLAD, please refer to the URL
2 papers · 0 benchmarks
InLegalNER is a corpus of 46545 annotated legal named entities mapped to 14 legal entity types.
2 papers · 1 benchmark
InSpaceType (Indoor Space Type Dataset for Monocular Depth Analysis)
High Quality Indoor Monocular Depth Estimation Dataset with focus on performance variation across space type - 1260 high quality evaluation pairs - Detailed inference variance across space types - Additional indoor image and depth pairs…
2 papers · 0 benchmarks
InVar-100 (Industrial Objects in Varied Contexts)
The Industrial Objects in Varied Contexts (InVar) Dataset was internally produced by our team and contains 100 objects in 20800 total images (208 images per class).
2 papers · 0 benchmarks
Inception Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
IndiaPoliceEvents is a corpus of 21,391 sentences from 1,257 English-language Times of India articles about events in the state of Gujarat during March 2002.
2 papers · 0 benchmarks
Indigo Mobile is a public dataset of copy detection patterns (CDP) based on DataMatrix modulation.
2 papers · 0 benchmarks
The Industrial Biscuits (Cookie) dataset is our internal dataset designed for the anomaly detection task, which captures Tarallini biscuits.
2 papers · 0 benchmarks
InfiniteBench (∞Bench: Extending Long Context Evaluation Beyond 100K Tokens)
Introduction Welcome to InfiniteBench, a cutting-edge benchmark tailored for evaluating the capabilities of language models to process, understand, and reason over super long contexts (100k+ tokens).
2 papers · 0 benchmarks
The Insider Threat Test Dataset is a collection of synthetic insider threat test datasets that provide both background and malicious actor synthetic data.
2 papers · 1 benchmark
InspiRe (Inspiring and non-inspiring posts from Reddit)
We analyze social media posts to tease out what makes a post inspiring and what topics are inspiring.
2 papers · 0 benchmarks
Includes two datasets published for the detection of fake and automated accounts.
2 papers · 0 benchmarks
InstaOrder can be used to understand the geometrical relationships of instances in an image.
2 papers · 0 benchmarks
Instantiation is a dataset for the task of instantiation detection
2 papers · 0 benchmarks
A curated dataset using the methodology of the paper is available in the Dataset folder.
2 papers · 0 benchmarks
This dataset contains data collected from 54 sensors deployed in the Intel Berkeley Research lab between February 28th and April 5th, 2004.
2 papers · 0 benchmarks
IOT BENIGN AND ATTACK TRACES Data Collected for ACM SOSR 2019 Attack & Benign Data Instructions Flow data contains flow counters of MUD flow, each instance in the file are collected every one minute.
2 papers · 0 benchmarks
IOT TRAFFIC TRACES Data Collected for IEEE TMC 2018 Cite our data A.
2 papers · 0 benchmarks
JAMUL (JApanese MUlti-Length Headline Corpus)
A large-scale evaluation dataset for headlines of three different lengths composed by professional editors.
2 papers · 0 benchmarks
JEMMA is an Extensible Java Dataset for ML4Code Applications, which is a large-scale dataset targeted at ML4 code.
2 papers · 0 benchmarks
The Jejueo Interview Transcripts (JIT) dataset is a parallel corpus containing 170k+ Jejueo-Korean sentences.
2 papers · 0 benchmarks
JNC (Japanese News Corpus)
The JNC data provides common supervision data for headline generation.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.