Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 215 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10273–10320 of 12,172
SLATS is a dataset which covers two data domains.
1 paper · 0 benchmarks
It consists of 32x32 pixel images of shapes with multiple attributes (size, location, rotation, color).
1 paper · 0 benchmarks
A collection of images of simple food items with various ground-truths
1 paper · 0 benchmarks
SimpleStories is a dataset of >2 million model-generated short stories.
1 paper · 0 benchmarks
Photometric stereoscopic test data sets under six lights taken using laboratory equipment.
1 paper · 0 benchmarks
Electromagnetic (EM) showers simulated dataset.
1 paper · 0 benchmarks
The dataset consists of 90 000 grayscale videos that show two objects of equal shape and size in which one object approaches the other one.
1 paper · 0 benchmarks
Dataset Details This dataset is primarily created for the work Fast muon tracking with machine learning implemented in FPGA (Arxiv link) that contains ~3M simulated muon events with Geant4.
1 paper · 0 benchmarks
Simulated pulse Doppler radar signatures for four classes of helicopter-like targets.
1 paper · 0 benchmarks
FLORIS farm dataset A dataset for graph neural network modeling of wind farms.
1 paper · 0 benchmarks
SinGAN-Seg-polyps is a synthetic dataset for polyp segmentation consisting of 10,000 synthetic polyps and masks.
1 paper · 0 benchmarks
This data comprises processed weather, soil, yield, and cultivation area for corn yield prediction in Sub-Sahara Africa, with emphasis on Nigeria.
1 paper · 0 benchmarks
Over 4000 single cycle waveforms, 600 samples
1 paper · 0 benchmarks
A medium-scale synthetic 4D Light Field video dataset for depth (disparity) estimation.
1 paper · 0 benchmarks
Dataset of 374 photos of hand-drawn sketches of App Inventor apps used for development of the Sketch2aia model for automatic generation of App Inventor wireframes from hand-drawn sketches.
1 paper · 0 benchmarks
We present the first fine-grained dataset of 1,497 3D VR sketch and 3D shape pairs for 1,005 chair shapes with large shapes diversity from the ShapeNetCore dataset from 50 participants.
1 paper · 0 benchmarks
Skit-S2I (Skit-S2I: An Indian Accented Speech to Intent dataset)
This dataset for Intent classification from human speech covers 14 coarse-grained intents from the Banking domain.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset consists of virtual scenes rendered in MuJoCo with multiple views each presented in multiple modalities: image, and synthetic or natural language descriptions.
1 paper · 0 benchmarks
A comprehensive set of all Slovenian tweets posted in the 2018-2020 period, with retweet links and assigned hate speech classes.
1 paper · 0 benchmarks
We introduce a large-scale video dataset Slovo for Russian Sign Language task.
1 paper · 1 benchmark
This new dataset represents a subset of the ImageNet1k.
1 paper · 0 benchmarks
Small-Bench NLP is a benchmark for small efficient neural language models trained on a single GPU.
1 paper · 0 benchmarks
Smarty4covid (The smarty4covid dataset and knowledge base: a framework enabling interpretable analysis of audio signals)
Harnessing the power of Artificial Intelligence (AI) and m-health towards detecting new bio-markers indicative of the onset and progress of respiratory abnormalities/conditions has greatly attracted the scientific and research interest…
1 paper · 0 benchmarks
SmokEng is a dataset of 3144 tweets, which are selected based on the presence of colloquial slang related to smoking and analyze it based on the semantics of the tweet.
1 paper · 0 benchmarks
Data related to 1040 patients with Covid-19 admitted to hospitals in Iran have been collected.
1 paper · 0 benchmarks
A dataset of 4562 images of which 4152 images contain a soccer ball.
1 paper · 0 benchmarks
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset.
1 paper · 0 benchmarks
The SoccerTrack dataset comprises top-view and wide-view video footage annotated with bounding boxes.
1 paper · 0 benchmarks
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity 🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain Dataset Statistics | Category | Samples | Description |…
1 paper · 0 benchmarks
The SNS data (Valente et al., 2013) is a four-wave survey conducted in Los Angeles county, the United States, which features a sample of 1,795 high-school students.
1 paper · 0 benchmarks
The scene derives from photo-realistic HM3D datasets.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The scene derives from photo-realistic MP3D datasets.
1 paper · 0 benchmarks
Leaves from genetically unique Juglans regia plants were scanned using X-ray micro-computed tomography (microCT) on the X-ray μCT beamline (8.3.2) at the Advanced Light Source (ALS) in Lawrence Berkeley National Laboratory (LBNL),…
1 paper · 0 benchmarks
SolarDK is a dataset for the detection and localization of solar.
1 paper · 0 benchmarks
SoliDiffy Differencing Contract Pairs and Edit Scripts Dataset The project creates and maintains two main datasets to assist with research and evaluation of Solidity smart contract differencing: Mutated Contracts Dataset: The mutated…
1 paper · 0 benchmarks
A large-scale HVMT dataset named SoloDance.
1 paper · 0 benchmarks
Songdo Traffic (Songdo Traffic: High Accuracy Georeferenced Vehicle Trajectories from a Large-Scale Study in a Smart City)
The Songdo Traffic dataset delivers precisely georeferenced vehicle trajectories captured through high-altitude bird's-eye view (BeV) drone footage over Songdo International Business District, South Korea.
1 paper · 0 benchmarks
Songdo Vision (Songdo Vision: Vehicle Annotations from High-Altitude BeV Drone Imagery in a Smart City)
The Songdo Vision dataset provides high-resolution (4K, 3840×2160 pixels) RGB images annotated with categorized axis-aligned bounding boxes (BBs) for vehicle detection from a high-altitude bird’s-eye view (BeV) perspective.
1 paper · 1 benchmark
Sonicverse is a multisensory simulation platform with integrated audio-visual simulation for training household agents that can both see and hear.
1 paper · 0 benchmarks
A high quality sound prompted semantic segmentation dataset
1 paper · 0 benchmarks
We collect a dataset of 805 clean videos that show the action of pouring water in a container.
1 paper · 1 benchmark
arxiv : https://arxiv.org/abs/2304.11708 Accepted at 29th International Congress on Sound and Vibration (ICSV29).
1 paper · 1 benchmark
Ensemble Tagger Training and Testing Set This data includes two files: The training set used to create the SCANL Ensemble tagger [1] and the "unseen" testing set that includes words from systems that are not available in the training set.
1 paper · 0 benchmarks
A large image source dataset including 23 devices, 6978 nat images, 2427 flat original images, total 23,361 images, 38G size, with serial numbers and naming conventions standardized.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.