Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 143 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6817–6864 of 12,172
RoMQA is a benchmark for robust, multi-evidence, and multi-answer question answering (QA).
2 papers · 0 benchmarks
A special scene-graph for intelligent vehicles.
2 papers · 0 benchmarks
This dataset contains 200 famous songs in different genres (mostly in rock) and the beats and downbeat annotations are provided by T.
2 papers · 1 benchmark
Consists of real objects rolling on complex terrains (pool table, elliptical bowl, and random height-field).
2 papers · 0 benchmarks
We created a building-image paired dataset that contains more than 3K samples using our roof modeling tools.
2 papers · 0 benchmarks
Rosario Dataset (The Rosario dataset: Multisensor data for localization and mapping in agricultural environments)
Agricultural dataset collected on-board out weed removing robot.
2 papers · 0 benchmarks
RuMedBench is a benchmark dataset for Russian medical language understanding.
2 papers · 0 benchmarks
RuOpenBookQA is a QA dataset with multiple-choice elementary-level science questions which probe the understanding of core science facts.
2 papers · 1 benchmark
https://github.com/dialogue-evaluation/RuSentNE-evaluation
2 papers · 0 benchmarks
Includes Russian tweets and news comments from multiple sources, covering multiple stories, as well as text classification approaches to stance detection as benchmarks over this data in this language.
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
RuWorldTree is a QA dataset with multiple-choice elementary-level science questions, which evaluate the understanding of core science facts.
2 papers · 1 benchmark
RyanSpeech is a speech corpus for research on automated text-to-speech (TTS) systems.
2 papers · 0 benchmarks
Technical Information Dates range from 2017-09-11 to 2018-02-16 and the time interval is 1 minute.
2 papers · 0 benchmarks
S-TEST is a benchmark for measuring the specificity of the language of pre-trained language models.
2 papers · 0 benchmarks
S1SLC_CVDL (A COMPLEX-VALUED ANNOTATED SINGLE LOOK COMPLEX SENTINEL-1 SAR DATASET FOR COMPLEX-VALUED DEEP NETWORKS)
ABSTRACT Development of the Complex-Valued (CV) deep learning architectures has enabled us to exploit the amplitude and phase components of the CV Synthetic Aperture Radar (SAR) data.
2 papers · 0 benchmarks
The S2-100K dataset is a dataset of 100,000 multi-spectral satellite images and their corresponding locations (latitude / longitude coordinates of the image centroid) sampled from Sentinel-2 via the Microsoft Planetary Computer.
2 papers · 0 benchmarks
Our proposed Synthetic-to-Real benchmark for more practical visual DA (termed S2RDA) includes two challenging transfer tasks of S2RDA-49 and S2RDA-MS-39.
2 papers · 0 benchmarks
SA-Det-100k is a large-scale class-agnostic object detection dataset for Research Purposes only.
2 papers · 1 benchmark
SART is a collection of three datasets for Similarity, Analogies and Relatedness for the Tatar language.
2 papers · 0 benchmarks
A visual complexity dataset that compromises of more than 1,400 images from seven image categories relevant to the above research areas, namely Scenes, Advertisements, Visualization and infographics, Objects, Interior design, Art, and…
2 papers · 0 benchmarks
SB20 (Sugar Beet 2020 University of Bonn)
Video sequences captured at a field on Campus Kleinaltendorf (CKA), University of Bonn, captured by BonBot-I, an autonomous weeding robot.
2 papers · 0 benchmarks
Dataset of complexity metrics paired with vulnerable smart contracts.
2 papers · 0 benchmarks
Contains16 burst images using smartphones for burst/video denoising, restoration, and enhancement tasks.
2 papers · 0 benchmarks
This dataset contains 63 signed distance function shaders collected mostly from Shadertoy.
2 papers · 0 benchmarks
Dataset of 676 security vulnerabilities patches.
2 papers · 0 benchmarks
The SEED-VIG dataset is composed of four parts.
2 papers · 0 benchmarks
SEPE 8K dataset is made of 40 different 8K (8192 x 4320) video sequences and 40 variant 8K (8192 x 5464) images.
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
We introduce the SHAD3S dataset, that for a given contour representation of a mesh, under a given illumination condition, provides the illumination masks on the object, a shadow mask on the ground, its diffuse and sketch renders.
2 papers · 0 benchmarks
SI-SCORE is a synthetic dataset for the analysis of robustness to object location, rotation and size.
2 papers · 0 benchmarks
SIDAR is a dataset designed to be a training and evaluation set for a multitude of tasks involving image alignment and artifact removal, such as deep homography estimation, dense image matching, 2D bundle adjustment, inpainting, shadow…
2 papers · 0 benchmarks
SILVR (A Synthetic Immersive Large-Volume Plenoptic Dataset)
We present SILVR, a dataset of light field images for six-degrees-of-freedom navigation in large fully-immersive volumes.
2 papers · 0 benchmarks
SIMARA (SIMARA: a database for key-value information extraction from full-page handwritten documents)
Description We propose a new database for information extraction from historical handwritten documents.
2 papers · 2 benchmarks
This repository contains the SINGA:PURA dataset, a strongly-labelled polyphonic urban sound dataset with spatiotemporal context.
2 papers · 0 benchmarks
SK-VG is a dataset for Scene Knowledge-guided Visual Grounding, where the image content and referring expressions are not sufficient to ground the target objects, forcing the models to have a reasoning ability on the long-form scene…
2 papers · 0 benchmarks
SKILL-102 (SKILL 102 Lifelong Learning Dataset)
SKILL-102 consists of 102 image classification datasets.
2 papers · 0 benchmarks
Large-scale integration of photovoltaics (PV) into electricity grids is challenged by the intermittent nature of solar power.
2 papers · 0 benchmarks
SKSF-A consists of seven distinct styles drawn by professional artists.
2 papers · 1 benchmark
This dataset comprehends the 3D building information model (in IFC and Revit formats), manually elaborated based on the terrestrial laser scanner of the sequence 2 of ConSLAM, and the refined ground truth (GT) poses (in TUM format) of…
2 papers · 0 benchmarks
The dataset consists of source code and LLVM IR pairs generated from accepted and de-duped programming contest solutions.
2 papers · 0 benchmarks
Contents (As on March 4, 2019) -------- The text corpus contains running text from various free licensed sources.
2 papers · 0 benchmarks
This is an enhanced Market-1501 dataset labeled with SMPL annotations, ie 3D human shape and pose ground truth.
2 papers · 0 benchmarks
SNDZoo (The Softwarised Network Data Zoo)
The softwarised network data zoo (SNDZoo) is an open collection of software networking data sets aiming to streamline and ease machine learning research in the software networking domain.
2 papers · 0 benchmarks
https://github.com/rail-berkeley/soar?tab=readme-ov-file#using-soar-data
2 papers · 0 benchmarks
SPARF is a large-scale ShapeNet-based synthetic dataset for novel view synthesis consisting of ~17 million images rendered from nearly 40,000 shapes at high resolution (400×400 pixels).
2 papers · 0 benchmarks
Training data for Hebrew morphological word segmentation
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.