Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 117 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5569–5616 of 12,172
This dataset provides the VCIP 2020 Grand Challenge on the NIR Image Colorization dataset.
3 papers · 1 benchmark
NKL (short for NanKai Lines) is a dataset for semantic line detection.
3 papers · 1 benchmark
NOD (Night Object Detection)
This is a high-quality large-scale Night Object Detection (NOD) dataset of outdoor images targeting low-light object detection.
3 papers · 0 benchmarks
NarraSum is a large-scale narrative summarization dataset.
3 papers · 0 benchmarks
Natural Hazards is a natural disaster dataset with sentiment labels, which contains nearly 50,00 Twitter data about different natural disasters in the United States (e.g., a tornado in 2011, a hurricane named Sandy in 2012, a series of…
3 papers · 0 benchmarks
Replay data from human players and AI agents navigating in a 3D game environment.
3 papers · 0 benchmarks
The largest dataset of extracted visual content from historic newspapers ever produced.
3 papers · 0 benchmarks
Contains harmful questions across different topics
3 papers · 0 benchmarks
NinaPro DB2 (DB2 - 40 Intact Subjects - Delsys Trigno electrodes)
The second Ninapro database includes 40 intact subjects and it is thoroughly described in the paper: "Manfredo Atzori, Arjan Gijsberts, Claudio Castellini, Barbara Caputo, Anne-Gabrielle Mittaz Hager, Simone Elsig, Giorgio Giatsidis,…
3 papers · 0 benchmarks
NusaCrowd is a collaborative initiative to collect and unite existing resources for Indonesian languages, including opening access to previously non-public resources.
3 papers · 0 benchmarks
The dataset contains Amazon products from 10 product categories with full human annotations.
3 papers · 2 benchmarks
The OAB Exams dataset is a valuable resource used in the context of legal information systems.
3 papers · 1 benchmark
OAM-TCD is a dataset of around 5k aerial images from around the world to support robust tree detection algorithms.
3 papers · 0 benchmarks
The OCTAGON dataset is a set of Angiography by Octical Coherence Tomography images (OCT-A) used to the segmentation of the Foveal Avascular Zone (FAZ).
3 papers · 0 benchmarks
ODSQA (Open-Domain Spoken Question Answering)
The ODSQA dataset is a spoken dataset for question answering in Chinese.
3 papers · 0 benchmarks
OFDIW (OnFocus Detection In the Wild)
OnFocus Detection In the Wild (OFDIW) is an onfocus detection dataset.
3 papers · 0 benchmarks
OGTD (Offensive Greek Tweet Dataset)
A manually annotated dataset containing 4,779 posts from Twitter annotated as offensive and not offensive.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
The OOP benchmark features 431 Python programs that encompass essential OOP concepts and features like classes and encapsulation methods¹².
3 papers · 0 benchmarks
OQMD v1.2 (The Open Quantum Materials Database)
The OQMD is a database of DFT calculated thermodynamic and structural properties of one million materials, created in Chris Wolverton's group at Northwestern University.
3 papers · 1 benchmark
OREBA (Objectively Recognizing Eating Behavior and Associated Intake)
The OREBA dataset aims to provide a comprehensive multi-sensor recording of communal intake occasions for researchers interested in automatic detection of intake gestures.
3 papers · 0 benchmarks
Evaluate radar localization in diverse environments Download: https://drive.google.com/drive/folders/1uATfrAe-KHlz29e-Ul8qUbUKwPxBFIhP Download
3 papers · 0 benchmarks
ORVS (Online Retinal image for Vessel Segmentation (ORVS))
The ORVS dataset has been newly established as a collaboration between the computer science and visual-science departments at the University of Calgary.
3 papers · 0 benchmarks
Is one of the largest egocentric datasets in the object search task with eyetracking information available Source: Deep Future Gaze: Gaze Anticipation on Egocentric Videos Using Adversarial Networks
3 papers · 0 benchmarks
https://athinagroup.eng.uci.edu/projects/ovrseen/
3 papers · 0 benchmarks
This dataset contains a set of face images taken between April 1992 and April 1994 at AT&T Laboratories Cambridge.
3 papers · 1 benchmark
The OpeReid dataset is a person re-identification dataset that consists of 7,413 images of 200 persons.
3 papers · 0 benchmarks
OpenCHAIR is a benchmark for evaluating open-vocabulary hallucinations in image captioning models.
3 papers · 0 benchmarks
The OpenCitations Meta database stores and delivers bibliographic metadata for all publications involved in the OpenCitations Index.
3 papers · 0 benchmarks
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
3 papers · 1 benchmark
OpenRooms FF(Forward Facing) is a dataset that extends OpenRooms into a multi-view setup.
3 papers · 0 benchmarks
OSAI introduces OpenTTGames - an open dataset aimed at evaluation of different computer vision tasks in Table Tennis: ball detection, semantic segmentation of humans, table and scoreboard and fast in-game events spotting.
3 papers · 0 benchmarks
OpenTrench3D, the first publicly available point cloud dataset of underground utilities from open trenches.
3 papers · 1 benchmark
A benchmark designed to evaluate MLLMs’ proficiency in understanding inter-object relationships and textual content.
3 papers · 0 benchmarks
P3 (Psychophysical Patterns Dataset)
A set of patterns used in psychophysical research to evaluate the ability of saliency algorithms to find targets distinct from distractors in orientation, color and size.
3 papers · 0 benchmarks
PAC (Polish Abusive Clauses)
''I have read and agree to the terms and conditions'' is one of the biggest lies on the Internet.
3 papers · 0 benchmarks
PARADE contains paraphrases that overlap very little at the lexical and syntactic level but are semantically equivalent based on computer science domain knowledge, as well as non-paraphrases that overlap greatly at the lexical and…
3 papers · 0 benchmarks
PCD (Poem Comprehensive Dataset)
The Arabic dataset is scraped mainly from الموسوعة الشعرية and الديوان.
3 papers · 3 benchmarks
PDE dataset (Parametric Partial Differential Equation dataset)
Contains data of parametric PDEs - Burgers' equation - Darcy's flow - Navier-Stokes equation
3 papers · 0 benchmarks
PDEs (Some PDE solutions)
In this dataset, you will find solutions of the following partial differential equations: - Burgers - Kortweg-de-Vries -Newell-Whitehead - Kuramoto-Sivashinsky You will find more info about how these were generated in the supplementary…
3 papers · 0 benchmarks
PDFVQA: A New Dataset for Real-World VQA on PDF Documents
3 papers · 0 benchmarks
The original dataset from Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting contains 6 months of traffic readings from 01/01/2017 to 05/31/2017 collected every 5 minutes by 325 traffic sensors in San…
3 papers · 1 benchmark
PETRAW (PEg TRAnsfer Workflow recognition by different modalities)
PETRAW data set was composed of 150 sequences of peg transfer training sessions.
3 papers · 6 benchmarks
PIC (Person In Context 2021)
The Person In Context (PIC) dataset is a dataset for human-centric relation segmentation (HRS), which contains 17,122 high-resolution images and densely annotated entity segmentation and relations, including 141 object categories, 23…
3 papers · 0 benchmarks
PNT (Parsing Time Normalizations)
The Parsing Time Normalizations (PNT) corpus in SCATE format allows the representation of a wider variety of time expressions than previous approaches.
3 papers · 1 benchmark
POPGym (Partially Observable Process Gym)
POPGym is designed to benchmark memory in deep reinforcement learning.
3 papers · 0 benchmarks
A large-scale video portrait dataset that contains 291 videos from 23 conference scenes with 14K frames.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.