Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 185 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8833–8880 of 12,172
This dataset contains images taken from camera traps set up in the Jura and Ain counties in France.
1 paper · 0 benchmarks
This dataset consists of ~350k JPEG images of streetlight columns installed on a public road infrastructure located in the city of Bristol, UK.
1 paper · 0 benchmarks
Image网 (pronounced Imagewang; 网 means "net" in Chinese) is an image classification dataset combined from Imagenette and Imagewoof datasets in a way to make it into a semi-supervised unbalanced classification problem: the validation set is…
1 paper · 0 benchmarks
ImagiFilter focusses on photographic and/or natural images, a very common use-case in computer vision research.
1 paper · 0 benchmarks
Imbalanced-MiniKinetics200 was proposed by "Minority-Oriented Vicinity Expansion with Attentive Aggregation for Video Long-Tailed Recognition" to evaluate varying scenarios of video long-tailed recognition.
1 paper · 0 benchmarks
Consists of 10,000 images is constructed, in which all the immediacy measures and the human poses are annotated.
1 paper · 0 benchmarks
This dataset comprises three immobilized fluorescently stained zebrafish imaged through the eXtended Field of view Light Field Microscope (XLFM, also known as Fourier Light Field Microscope).
1 paper · 0 benchmarks
As a first step towards building models that can recognise immune cells in WSIs, we introduce Immunocto, a high-resolution (40 x magnification) massive database of 2,310,257 immune cells distributed across 4 immune cell subtypes (CD4…
1 paper · 0 benchmarks
This dataset contains 6,387 ChatGPT prompts collected from four platforms (Reddit, Discord, websites, and open-source datasets) during Dec 2022 to May 2023.
1 paper · 0 benchmarks
About A realistic visual-inertial dataset with 58 sequences spanning 5km of trajectories and 1.5 hours of recordings, designed for evaluating SLAM systems in indoor pedestrian-rich environments.
1 paper · 0 benchmarks
AI algorithms, and in particular Machine Learning (ML) algorithms, learn from data tasks that have been traditionally done by humans such as: image classification, facial recognition, linguistic translation etc.
1 paper · 0 benchmarks
InHARD (Industrial Human Action Recognition Dataset in the Context of Industrial Collaborative Robotics)
We introduce a RGB+S dataset named “Industrial Human Action Recognition Dataset” (InHARD) from a real-world setting for industrial human action recognition with over 2 million frames, collected from 16 distinct subjects.
1 paper · 0 benchmarks
Simulates unanticipated user needs in the deployment stage.
1 paper · 0 benchmarks
There was no predefined dataset of party symbols to be usedas a benchmark.
1 paper · 0 benchmarks
We present two multi-modal datasets, one for Main Board IPOs, and the other for Small and Medium Enterprises (SME) IPOs.
1 paper · 0 benchmarks
Indiscapes2, a new large-scale diverse dataset of Indic manuscripts with semantic layout annotations.
1 paper · 0 benchmarks
The Deepfake face detection task involves a facial image of unknown authenticity for testing.
1 paper · 0 benchmarks
The dfdindoor dataset contains 110 images for training and 29 images for testing.
1 paper · 0 benchmarks
http://redwood-data.org/indoorlidarrgbd/
1 paper · 0 benchmarks
IndraEye (IndraEye: Infrared Electro-Optical Drone-based Aerial Object Detection Dataset)
Deep neural networks (DNNs) have demonstrated superior performance when trained on well-illuminated environments, given that the images are captured through an Electro-Optical (EO) camera, which offers rich texture content.
1 paper · 0 benchmarks
This is a real-world industrial benchmark dataset from a major medical device manufacturer for the prediction of customer escalations.
1 paper · 0 benchmarks
A dataset of books for very young children.
1 paper · 0 benchmarks
A collection of large languge model responses to tasks of propositional logic.
1 paper · 0 benchmarks
The goal of InfoLossQA is to generate a series of QA pairs that reveal to lay readers what information a simplified text lacks compared to its original.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
InfraParis is a novel and versatile dataset supporting multiple tasks across three modalities: RGB, depth, and infrared.
1 paper · 0 benchmarks
This is a dataset that catalogs 2.6 million patents granted between 2005 and 2017.
1 paper · 0 benchmarks
Inshorts News dataset Inshorts provides a news summary in 60 words or less.
1 paper · 1 benchmark
InstaCities1M is a dataset of social media images with associated text.
1 paper · 0 benchmarks
InstructOpenWiki is a substantial instruction tuning dataset for Open-world IE enriched with a comprehensive corpus, extensive annotations, and diverse instructions.
1 paper · 0 benchmarks
A dataset for image editing containing >450k samples of: 1.
1 paper · 0 benchmarks
The training subset consists of 15 robotic nephrectomy procedures captured on the da Vinci X or Xi system.
1 paper · 0 benchmarks
This newly curated synthetic dataset specifies an additional reference region to guide image harmonization.
1 paper · 0 benchmarks
A benchmark for visual intuitive physics reasoning.
1 paper · 0 benchmarks
Semantic segmentation of drone images is critical for various aerial vision tasks as it provides essential semantic details to understand scenes on the ground.
1 paper · 0 benchmarks
For the purpose of training and evaluating our intent classification model for electric automation, we curated a dataset consisting of intent-based user instructions.
1 paper · 0 benchmarks
This dataset is derived from the Waymo Motion dataset and focuses on capturing the interactions between autonomous vehicles (AVs) and traffic control devices such as traffic lights and stop signs.
1 paper · 0 benchmarks
Interactive Gibson is a comprehensive benchmark for training and evaluating Interactive Navigation: robot navigation strategies where physical interaction with objects is allowed and even encouraged to accomplish a task.
1 paper · 0 benchmarks
The dataset contains summary statistics and engagement metrics captured from users in a live, 'in-the-wild' study of an interactive TV show.
1 paper · 0 benchmarks
A dataset of sentence pairs annotated following the formalization.
1 paper · 0 benchmarks
The dataset concerns ADL activities performed in a smart home environment in an Interwoven manner.
1 paper · 0 benchmarks
Invisible Mobile Keyboard Dataset contains user initial, age, type of mobile devices, size of the screen, time taken for typing each phrase, and annotation of typed phrases with coordinate values of the typed position (x and y points).
1 paper · 0 benchmarks
ABSTRACT Recently, the technology of the fourth revolution has given the characteristics of things constantly expanding, and everything, including people, things, people, and the environment, is connected based on the Internet.
1 paper · 0 benchmarks
IoT-23 (IoT-23: A labeled dataset with malicious and benign IoT network traffic)
IoT-23 is a dataset of network traffic from Internet of Things (IoT) devices.
1 paper · 0 benchmarks
https://github.com/uci-plrg/iotcheck-data
1 paper · 0 benchmarks
The dataset includes source code vulnerabilities in some of the most commonly used IoT frameworks.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.