Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 216 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10321–10368 of 12,172
Source code (Source code underlying the publication: Topology-Based Reconstruction Prevention for Decentralised Learning)
MATLAB code to reproduce results presented in the paper "Topology-Based Reconstruction Prevention for Decentralised Learning".
1 paper · 0 benchmarks
Scene Graph Generation (SGG) converts visual scenes into structured graph representations, providing deeper scene understanding for complex vision tasks.
1 paper · 0 benchmarks
The Spaceship dataset is a dataset for evaluating agents’ ability to learn to solve a class of physics-based tasks.
1 paper · 0 benchmarks
SpamHunter (Spam SMS dataset collected from Twitter spam reports.)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Classifying Email as Spam or Non-Spam.
1 paper · 0 benchmarks
Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehension (MRC) or natural language inference (NLI).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Dataset built from partial reconstructions of real-world indoor scenes using RGB-D sequences from ScanNet, aimed at estimating the unknown position of an object (e.g.
1 paper · 0 benchmarks
Insects are the most important global pollinator of crops and play a key role in maintaining the sustainability of natural ecosystems.
1 paper · 0 benchmarks
SpeakGer (SpeakGer: A meta-data enriched speech corpus of German state and federal parliaments)
A dataset of German parliament debates covering 74 years of plenary protocols across all 16 state parliaments of Germany as well as the German Bundestag.
1 paper · 0 benchmarks
Spectral Detection and Analysis Based Paper(SDAAP) dataset is the first open-source textual knowledge dataset for spectral analysis and detection and contains annotated literature data as well as corresponding knowledge instruction data,…
1 paper · 0 benchmarks
SpectroVision is a dataset of 14,400 high resolution texture images and spectral measurements collected from a PR2 mobile manipulator that interacted with 144 household objects from eight material categories.
1 paper · 0 benchmarks
Dataset Summary Speech Brown is a comprehensive, synthetic, and diverse paired speech-text dataset in 15 categories, covering a wide range of topics from fiction to religion.
1 paper · 0 benchmarks
A high quality speech prompted semantic segmentation dataset
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Spiced is a paraphrase dataset of scientific findings annotated for degree of information change.
1 paper · 0 benchmarks
Synthetic soccer players rendered on top of real world stadium images in 4K covering half a pitch each.
1 paper · 1 benchmark
Overview The Spike-X4K Dataset is a high-resolution image reconstruction resource tailored for the latest advancements in spike camera technology.
1 paper · 1 benchmark
Spot the Difference Corpus is a corpus of task-oriented spontaneous dialogues which contains 54 interactions between pairs of subjects interacting to find differences in two very similar scenes.
1 paper · 0 benchmarks
This is the synthetic dataset that is introduced in the paper https://arxiv.org/abs/2403.03375.
1 paper · 0 benchmarks
This dataset contains over 47,000 LEGO structures of over 28,000 unique 3D objects accompanied by detailed captions.
1 paper · 0 benchmarks
The dataset identifies the shortcomings of existing benchmarks in evaluating the problem of compositional generalization, which underscores the need for the development of datasets tailored to assess compositional generalization in open…
1 paper · 1 benchmark
The dataset contains 256x256 tiles extracted from Whole Slide Images (WSI) of mouse liver tissue stained with H&E and Masson's Trichrome.
1 paper · 0 benchmarks
The datasets used in our WACV paper High-Fidelity Document Stain Removal via A Large-Scale Real-World Dataset and A Memory-Augmented Transformer.
1 paper · 0 benchmarks
3D confocal stacks with corresponding 2D Light-field microscope images Confocal: -Single volume dimension: 1287x1287x64.
1 paper · 0 benchmarks
Dataset of NP-Hard Job Shop Scheduling Problem (JSSP), specifically designed for LLM fine-tuning.
1 paper · 0 benchmarks
This data contains throughput, RTT, power consumption and speed data measured with Starlink during mobility setups.
1 paper · 0 benchmarks
> The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents > > Xing Han Lu, Siva Reddy, Harm de Vries > > EACL 2023 | | | | | | | :--: | :--: | :--: | :--: | :--: | | Code | Huggingface | Request on…
1 paper · 1 benchmark
When arriving at each state, each observation token gets a coin toss to see whether it will appear in the output observation string.
1 paper · 0 benchmarks
The instances were drawn randomly from a database of 7 outdoor images.
1 paper · 0 benchmarks
Energy consumption data is collected using IoT based systems and used for prediction.
1 paper · 0 benchmarks
Steerability probe example for text-rewriting.
1 paper · 0 benchmarks
Steredo Waterdrop is a real-world dataset for research on stereo waterdrop removal.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Stickers is a dataset consisting of 577 high-quality sticker images with alpha channel.
1 paper · 0 benchmarks
The Store Dataset is a dataset for estimating 3D poses of multiple humans in real-time.
1 paper · 0 benchmarks
StoryDB is a broad multi-language dataset of narratives.
1 paper · 0 benchmarks
Visual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations.
1 paper · 0 benchmarks
This dataset is a "part I" extension of the "Engineered cardiac microbundle time-lapse microscopy image dataset" and contains 732 experimental time-lapse image sequences of beating hiPSC-based cardiac microbundles using microbundle strain…
1 paper · 0 benchmarks
A real-world image dataset that contains more than 900 images generated from 26 street cameras and 7 object categories annotated with detailed bounding box.
1 paper · 0 benchmarks
A large-scale dataset composed of object-centric street view scenes along with point correspondences and camera pose information.
1 paper · 0 benchmarks
Streetscore (STREETSCORE--PREDICTING THE PERCEIVED SAFETY OF ONE MILLION STREETSCAPES)
Paper abstract: Social science literature has shown a strong connection between the visual appearance of a city’s neighborhoods and the behavior and health of its citizens.
1 paper · 0 benchmarks
Student Essay is widely used in research on argument segmentation
1 paper · 0 benchmarks
In this Coding and Coding Scheme spreadsheet: Student answers to reflection questions from pre-context workshops; coding scheme for student reflections; and coding of student reflection
1 paper · 0 benchmarks
This dataset consists of EEG (Electroencephalogram) recordings collected from students at our college during an educational experiment.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.