Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 201 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9601–9648 of 12,172
Dataset provided along with the repository of the paper: https://github.com/hardware-fab/DL-to-locate-COs-for-SCA/tree/main
1 paper · 0 benchmarks
The Swiss Drone data set was recorded around Cheseaux-sur-Lausanne in Switzerland using a senseFly eBee Classic in 2014 (SenseFly, 2020).
1 paper · 1 benchmark
Large Language Models (LLMs) have the potential to enhance Agent-Based Modeling by better representing complex interdependent cybersecurity systems, improving cybersecurity threat modeling and risk management.
1 paper · 0 benchmarks
Olympic 2024 is a human-annotated dataset that contains 220 high-quality instance.
1 paper · 0 benchmarks
Omiverse Object is a large-scale synthetic dataset of 60,000 images including both transparent and opaque objects in different scenes.
1 paper · 0 benchmarks
Omni-Image is built as a challenging but tractable dataset for continual learning and few-shot learning.
1 paper · 0 benchmarks
The Omni-MOT is realistic CARLA based large-scale dataset with over 14M frames for multiple vehicle tracking .
1 paper · 0 benchmarks
In order to evaluate the effectiveness of NToP in real-world scenarios, we collect a new dataset OmniLab with a top-view omnidirectional camera, mounted on the ceiling of two different rooms (bedroom, living room) at 2.5 m height.
1 paper · 0 benchmarks
To effectively evaluate OmniCount across open-vocabulary, supervised, and few-shot counting tasks, a dataset catering to a broad spectrum of visual categories and instances featuring various visual categories with multiple instances and…
1 paper · 2 benchmarks
Omnipush is a dataset with high variety of planar pushing behavior.
1 paper · 0 benchmarks
A dataset for online novel recommendation.
1 paper · 0 benchmarks
The OnlySports Dataset is a comprehensive collection of sports-related text data, comprising approximately 600 billion tokens.
1 paper · 0 benchmarks
OntoRock is a benchmark for evaluating the robustness of existing NER models via a systematic evaluation protocol.
1 paper · 0 benchmarks
A biomedical dataset supporting ontology enrichment from texts, by concept discovery and placement, adapting the MedMentions dataset (PubMed abstracts) with SNOMED CT of versions in 2014 and 2017 under the Diseases (disorder) sub-category…
1 paper · 0 benchmarks
Open MIC (Open Museum Identification Challenge)
Open MIC (Open Museum Identification Challenge) contains photos of exhibits captured in 10 distinct exhibition spaces of several museums which showcase paintings, timepieces, sculptures, glassware, relics, science exhibits, natural history…
1 paper · 0 benchmarks
Dataset of cross-layer Radio Access Network (RAN) Key Performance Measurements (KPMs) and protocol stack logs collected on an Open RAN deployment instantiated on Colosseum with traffic twinned from that of commercial cellular traces.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
🏃♂️ Open-HypermotionX Dataset Open-Hypermotion is a large-scale, high-quality dataset designed for training and evaluating pose-guided human image animation models, with a special focus on complex, dynamic human motions (Hypermotion),…
1 paper · 0 benchmarks
A human-refined dataset of OpenAPI definitions based on the APIs.guru OpenAPI directory.
1 paper · 1 benchmark
OpenAlex (OpenAlex: The open catalog to the global research system)
Details of OpenAlex Data can be seen in the official wensite: https://openalex.org/about
1 paper · 0 benchmarks
OpenD5 is a a meta-dataset which aggregates 675 open-ended problems ranging across business, social sciences, humanities, machine learning, and health, and uses a set of unified evaluation metrics: validity, relevance, novelty, and…
1 paper · 0 benchmarks
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
1 paper · 0 benchmarks
OpenGDA is a benchmark for evaluating graph domain adaptation models.
1 paper · 0 benchmarks
OpenING is a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks.
1 paper · 0 benchmarks
OpenREACT-CHON-EFH (OpenREACT-CHON-EFH — Open REaction Dataset of Atomic ConfiguraTions comprising C, H, O, N with Energies, Forces, and Hessians)
RTP Dataset (Reactant–Transition State–Product Dataset) The RTP dataset forms the core training and evaluation set and consists of 35,087 molecular geometries sampled from 11,961 unique elementary reactions.
1 paper · 0 benchmarks
We create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triples.
1 paper · 1 benchmark
OpenSpeaks Voice: Odia is a large speech dataset in the Odia language of India that is stewarded by Subhashish Panigrahi and is hosted at the O Foundation.
1 paper · 0 benchmarks
A high-resolution multi-sensor remote sensing scene classification dataset, appropriate for training and evaluating image classification models in the remote sensing domain.
1 paper · 0 benchmarks
OpenWPM Crawls is a dataset of 103 online, mostly mainstream news websites.
1 paper · 0 benchmarks
A large-scale dataset of measurements of ETSI ITS-G5 Dedicated Short Range Communications (DSRC) is presented.
1 paper · 0 benchmarks
This repository contains the datasets corresponding to the two benchmark problems appearing in the SIAM papers "The Random Feature Model for Input-Output Maps between Banach Spaces" [SIAM J.
1 paper · 0 benchmarks
URL:https://huggingface.co/datasets/Aurora-Gem/OptMATH-Train
1 paper · 0 benchmarks
Image corruptions modelling primary optical aberrations.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Orchard (A Benchmark For Measuring Systematic Generalization of Multi-Hierarchical Reasoning)
Orchard is a diagnostic dataset for systematically evaluating hierarchical reasoning in state-of-the-art neural sequence models
1 paper · 0 benchmarks
Orchid2024 is a fine-grained classification dataset specifically designed for Chinese Cymbidium orchid cultivars.
1 paper · 0 benchmarks
It includes 10 data sets that consists of both raw data set and encoded data set where it is encoded through BERT-Sort Encoder with MLM initialization of .
1 paper · 1 benchmark
This Dataset contains pairs off textual natural language questions and SPARQL queries on a small organizational graph(https://github.com/AKSW/AI-Tomorrow-2023-KG-ChatGPT-Experiments/blob/main/FoafVcardOrg/foaf-vcard-org-data.ttl) which was…
1 paper · 0 benchmarks
The Out the Window (OTW) dataset is a crowdsourced activity dataset containing 5,668 instances of 17 activities from the NIST Activities in Extended Video (ActEV) challenge.
1 paper · 0 benchmarks
Monitoring and evaluating of driving behavior is the main goal of this paper that encourage us to develop a new system based on Inertial Measurement Unit (IMU) sensors of smartphones.
1 paper · 0 benchmarks
Overnight is a dataset for semantic parsing in eight domains.
1 paper · 0 benchmarks
The ontology files, readme and statistical information can be found and browsed in the ontology library.
1 paper · 0 benchmarks
The Oxford Road Boundaries is a dataset designed for training and testing machine-learning-based road-boundary detection and inference approaches.
1 paper · 0 benchmarks
The Oxford Town Center dataset is a 5-minute video with 7500 frames annotated, which is divided into 6500 for training and 1000 for testing data for pedestrian detection.
1 paper · 1 benchmark
P-OCT (Peripapillary OCT Images)
The entire dataset consists of 61 different subjects, for each of which 12 radial OCT B-scans are collected at the Ophthalmology Department of Shanghai General Hospital by using DRI OCT-1 Atlantis (Topcon Corporation, Tokyo, Japan).
1 paper · 0 benchmarks
P4D prompts (P4D universal jailbreaking prompt for T2I models)
This dataset contains prompts designed to evaluate and challenge the safety mechanisms of generative text-to-image models, with a particular focus on identifying prompts that are likely to produce images containing nudity.
1 paper · 0 benchmarks
This is the set of graphs used in the PACE 2022 challenge for computing the Directed Feedback Vertex Set, from the Heuristic track.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.