Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 248 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11857–11904 of 12,172
The dataset consists of 1024 x 1024 bitmap (.bmp) images, each containing a 16 x 16 array of image patches.
0 papers · 0 benchmarks
PhyMER (PhyMER: Physiological Dataset for Multimodal Emotion Recognition With Personality as a Context)
A physiological signal dataset with multiple physiological signals collected from 30 participants.
0 papers · 0 benchmarks
Physionet MI dataset: https://physionet.org/pn4/eegmmidb/ This data set consists of over 1500 one- and two-minute EEG recordings, obtained from 109 volunteers [2].
0 papers · 0 benchmarks
There are about 208 000 jokes in this database scraped from three sources.
0 papers · 0 benchmarks
Plant Centroids is a dataset for stem emerging points (SEP) detection in RGB and NIR image data.
0 papers · 0 benchmarks
PolandNFC is a collection of 4,000 recordings from Hanna Pamuła's PhD project of monitoring autumn nocturnal bird migration.
0 papers · 0 benchmarks
Pose Estimation Lunar Robot (Dataset for camera pose estimation research using computer simulated images from rovers on the lunar surface)
Overview The goal: using simulation data to train neural networks to estimate the pose of a rover's camera with respect to a known target object The mission context: A simulated lunar surface, with lunar landers and lunar rovers.
0 papers · 0 benchmarks
Cancer genomics and precision oncology: The TCGA Research Network started in 2005 has profiled and analyzed a large number of human tumors to discover molecular aberrations at the DNA, RNA, protein, and epigenetic levels and thereby…
0 papers · 0 benchmarks
Dataset: RGB-D Images for Real-World and Synthetic Object Scenes This dataset consists of both real-world and synthetic RGB-D images, designed for object detection, classification, and segmentation tasks, particularly for primitive shape…
0 papers · 0 benchmarks
The Princeton Shape dataset provides a repository of 3D models and software tools for evaluating shape-based retrieval and analysis algorithms.
0 papers · 0 benchmarks
The PropBankPT (Branco et al., 2012) is a set of sentences annotated with their constituency structure and semantic role tags, composed of 3,406 sentences and 44,598 tokens taken from the Wall Street Journal translated.
0 papers · 0 benchmarks
QBQTC (QQ Browser Query Title Corpus)
The QQ Browser Query Title Corpus (QBQTC) is a large-scale dataset constructed for search scenarios by the QQ Browser search engine.
0 papers · 0 benchmarks
QPT (Quantum Process Tomography)
Quantum process tomography (QPT) is a method for experimentally reconstructing the quantum channel from measurement data.
0 papers · 0 benchmarks
QuAIL (Question Answering for Artificial Intelligence)
A new kind of question-answering dataset that combines commonsense, text-based, and unanswerable questions, balanced for different genres and reasoning types.
0 papers · 0 benchmarks
RAF-ML (Real-world Affective Faces Multi Label)
Real-world Affective Faces Multi Label (RAF-ML) is a multi-label facial expression dataset with around 5K great-diverse facial images downloaded from the Internet with blended emotions and variability in subjects' identity, head poses,…
0 papers · 0 benchmarks
Laser powder bed fusion (LBPF) is the additive manufacturing (3D printing) process for metals.
0 papers · 0 benchmarks
RF100-VL is a multi-domain benchmark for object detection.
0 papers · 0 benchmarks
The dataset contains 300 objects organized into 51 categories and has been made publicly available to the research community so as to enable rapid progress based on this promising technology.
0 papers · 0 benchmarks
RNA-Puzzles is a collective experiment for blind RNA structure prediction.
0 papers · 0 benchmarks
The preview of the road surface states is essential for improving the safety and the ride comfort of autonomous vehicles.
0 papers · 0 benchmarks
RSPECT (The RSNA Pulmonary Embolism CT)
The RSNA Pulmonary Embolism CT (RSPECT) Dataset is composed of CT pulmonary angiogram images and annotations related to pulmonary embolism.
0 papers · 0 benchmarks
RS_NS92 (Remote Sensing Natural Scenes 92 (classes))
Consists of 36,785 images belonging to a diverse 92 classes.
0 papers · 0 benchmarks
RTI International (RTI) generated 2,611 labeled point locations representing 19 different land cover types, clustered in 5 distinct agroecological zones within Rwanda.
0 papers · 0 benchmarks
The dataset has railway track images of two types: normal and defective The task is to classify a given image into normal or defective.
0 papers · 1 benchmark
A review on raw subjective scores and data manipulation for before and after refining Mean opinion Scores
0 papers · 0 benchmarks
This dataset consists of two categories.
0 papers · 0 benchmarks
Data Set Information: A data set describing the evolution of results in the Portuguese Parliamentary Elections of October 6th 2019.
0 papers · 0 benchmarks
RealWorldQA is a benchmark designed to evaluate the real-world spatial understanding capabilities of multimodal AI models.
0 papers · 1 benchmark
The Redteaming Resistance Benchmark is a project aimed at evaluating the robustness of language models, both open-source and black-box, through redteaming attacks.
0 papers · 0 benchmarks
The RefoMB dataset is part of a project called RLAIF-V, which stands for "Aligning MLLMs through Open-Source AI Feedback for Super GPT-4V Trustworthiness." It's an open-source multimodal preference dataset that contains more than 30,000…
0 papers · 0 benchmarks
Welcome to "Reliable Air Ambulance Services in Hyderabad: Your Guide to Emergency Medical Transport." This forum is dedicated to providing comprehensive information, support, and resources about air ambulance services in Hyderabad.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.