Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 139 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6625–6672 of 12,172
NeuroVoz (NeuroVoz: a Castillian Spanish corpus of parkinsonian speech)
The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis.
2 papers · 0 benchmarks
This dataset is recreated using offline augmentation from the original dataset.
2 papers · 2 benchmarks
A text corpus of more than 200,000 sentences from eleven news sources regarding Donald Trump.
2 papers · 0 benchmarks
News Interactions on Globo.com (News Portal User Interactions by Globo.com - A large dataset for news recommendations offline evaluation and analytics)
Context This large dataset with users interactions logs (page views) from a news portal was kindly provided by [Globo.com][1], the most popular news portal in Brazil, for reproducibility of the experiments with CHAMELEON - a…
2 papers · 0 benchmarks
News SEO Dataset (Detection and Discovery of Misinformation Sources using Attributed Webgraphs)
Search Engine Optimization (SEO) attributes provide strong signals for predicting news site reliability.
2 papers · 0 benchmarks
NewsPH-NLI is a sentence entailment benchmark dataset in the low-resource Filipino language.
2 papers · 0 benchmarks
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
2 papers · 0 benchmarks
NorDial is the first step to creating a corpus of dialectal variation of written Norwegian.
2 papers · 0 benchmarks
Abstract The Norwegian Endurance Athlete ECG Database contains 12-lead ECG recordings from 28 elite athletes from various sports in Norway.
2 papers · 0 benchmarks
Scene-focused, multi-modal, episodic data of the images and symbolic world-states seen by an agent completing a pogo-stick assembly task within a video game world.
2 papers · 0 benchmarks
The Online Action Detection Dataset (OAD) was captured using the Kinect V2 sensor, which collects color images, depth images and human skeleton joints synchronously.
2 papers · 1 benchmark
OADAT (OADAT: Experimental and Synthetic Clinical Optoacoustic Data for Standardized Image Processing)
An experimental and synthetic (simulated) OA raw signals and reconstructed image domain datasets rendered with different experimental parameters and tomographic acquisition geometries.
2 papers · 0 benchmarks
The dataset contains images of 16 artworks included in the cultural site “Galleria Regionale di Palazzo Bellomo2”.
2 papers · 1 benchmark
OCR-IDL (OCR Annotations for Industry Document Library Dataset)
The OCR-IDL dataset comprises the OCR annotations for a subset of 26M pages of the large-scale IDL document library.
2 papers · 0 benchmarks
The OCR-VQA dataset is a valuable resource for research in the field of Visual Question Answering (VQA).
2 papers · 0 benchmarks
A dataset containing the results of a MUSHRA listening test conducted with expert listeners from 2 international laboratories.
2 papers · 1 benchmark
ODMS (Object Depth via Motion and Segmentation)
ODMS is a dataset for learning Object Depth via Motion and Segmentation.
2 papers · 0 benchmarks
OIR is a financial-domain dataset of the outbound intent recognition task.
2 papers · 0 benchmarks
OLKAVS (An Open Large-Scale Korean Audio-Visual Speech Dataset)
The dataset contains 1,150 hours of transcribed audio from 1,107 Korean speakers in a studio setup with nine different viewpoints and various noise situations.
2 papers · 0 benchmarks
OPT (Object Pose Tracking)
Accurately tracking the six degree-of-freedom pose of an object in real scenes is an important task in computer vision and augmented reality with numerous applications.
2 papers · 1 benchmark
This is a large-scale dataset of quantum-mechanically calculated properties (DFT level) of crystalline materials for graph representation learning that contains approximately 900k entries (OQM9HK).
2 papers · 1 benchmark
ORCAS-I (Queries Annotated with Intent using Weak Supervision)
A labelled version of the ORCAS click-based dataset of Web queries, which provides 18 million connections to 10 million distinct queries.
2 papers · 1 benchmark
A new video dataset for OR, with 30, 000 objects over 5, 000 stereo video sequences annotated for their descriptions and gaze.
2 papers · 0 benchmarks
Description OV dataset is the camera calibration dataset.
2 papers · 0 benchmarks
OpenWebText2 is an enhanced version of the original OpenWebTextCorpus.
2 papers · 1 benchmark
ObjectNet3D is a large scale database for 3D object recognition, named, that consists of 100 categories, 90,127 images, 201,888 objects in these images and 44,147 3D shapes.
2 papers · 0 benchmarks
Occ-Traj120 is a trajectory dataset that contains occupancy representations of different local-maps with associated trajectories.
2 papers · 0 benchmarks
Occluded COCO is automatically generated subset of COCO val dataset, collecting partially occluded objects for a large variety of categories in real images in a scalable manner, where target object is partially occluded but the…
2 papers · 1 benchmark
From Schaub, Michael T., et al.
2 papers · 0 benchmarks
A realistic, diverse, and challenging dataset for object detection on images.
2 papers · 1 benchmark
OmniCity is a dataset for omnipotent city understanding from multi-level and multi-view images.
2 papers · 0 benchmarks
This Dataset is described in Charting the Landscape of Online Cryptocurrency Manipulation.
2 papers · 0 benchmarks
This is a transnational data set which contains all the transactions occurring between 01/12/2010 and 09/12/2011 for a UK-based and registered non-store online retail.
2 papers · 0 benchmarks
Only Time Will Tell (Time-respecting and time-ignoring horizon of code review network at Microsoft)
Simulation results of time-respecting and time-ignoring horizon of code review network at Microsoft as JSON.
2 papers · 0 benchmarks
A classification dataset of radar spectrograms in i "ground surveillance" setting recorded with the Open Radar Initiative.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
OpenAsp Dataset OpenAsp is an Open Aspect-based Multi-Document Summarization dataset derived from DUC and MultiNews summarization datasets.
2 papers · 0 benchmarks
OpenBMAT (Open Broadcast Media Audio from TV)
Open Broadcast Media Audio from TV (OpenBMAT) is an open, annotated dataset for the task of music detection that contains over 27 hours of TV broadcast audio from 4 countries distributed over 1647 one-minute long excerpts.
2 papers · 0 benchmarks
(L)ifel(O)ng (R)obotic V(IS)ion (OpenLORIS) - Object Recognition Dataset (OpenLORIS-Object) is designed for accelerating the lifelong/continual/incremental learning research and application,currently focusing on improving the continuous…
2 papers · 0 benchmarks
OpenSLR (Open Speech and Language Resources)
OpenSLR is a repository of open speech and language resources, including large-scale transcribed audio corpora and related software.
2 papers · 1 benchmark
OpenStreetView-5M establishes a new open benchmark for geolocation by providing a large, open, and clean dataset.
2 papers · 1 benchmark
OpenViDial 2.0 is a larger-scale open-domain multi-modal dialogue dataset compared to the previous version OpenViDial 1.0.
2 papers · 1 benchmark
This is the dataset used in the PACE 2016 challenge, Track B, which was computing minimal Feedback Vertex Set.
2 papers · 0 benchmarks
This is the set of graphs used in the PACE 2022 challenge for computing the Directed Feedback Vertex Set, from the Exact track.
2 papers · 0 benchmarks
PAD Dataset (Pose-agnostic/Multi-pose Anomaly Detection Dataset)
Multi-pose Anomaly Detection (MAD) dataset, which represents the first attempt to evaluate the performance of pose-agnostic anomaly detection.
2 papers · 1 benchmark
PAL4Inpaint is a dataset consisting of 4,795 inpainting results with per-pixel perceptual artifacts annotations designed for image inpainting tasks.
2 papers · 0 benchmarks
Enables research on early detection of sexual predators in chats (eSPD).
2 papers · 0 benchmarks
Appearance-based gaze estimation systems have shown great progress recently, yet the performance of these techniques depend on the datasets used for training.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.