Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 250 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11953–12000 of 12,172
Given two sentences, the participants are asked to determine whether they express the same or very similar meaning and optionally a degree score between 0 and 1.
0 papers · 0 benchmarks
Semeion (Semeion Handwritten Digit Data Set)
1593 handwritten digits from around 80 persons were scanned, stretched in a rectangular box 16x16 in a gray scale of 256 values.
0 papers · 0 benchmarks
It is released by the Shanghai Central Meteorological Observatory (SCMO) in 2020, records serval years of historical precipitation events in the Yangtze River delta area.
0 papers · 0 benchmarks
The ShapeNoiseHorseBird dataset is a curated collection designed to challenge shape recognition models with varying levels of noise.
0 papers · 0 benchmarks
This dataset can be used as a benchmark for clustering word embeddings for German.
0 papers · 0 benchmarks
sheeep ruminate behavior image and label
0 papers · 0 benchmarks
ShipRSImageNet is a large-scale fine-grainted dataset for ship detection in high-resolution optical remote sensing images.
0 papers · 0 benchmarks
This dataset contains images and annotations for scene text detection and recognition.
0 papers · 0 benchmarks
SimNICT is the first dataset for training universal non-ideal measurement CT (NICT) enhancement models.
0 papers · 0 benchmarks
SimSceneTVB is a dataset of 600 simulated sound scenes of 45s each representing urban sound environments, simulated using the simScene Matlab library.
0 papers · 0 benchmarks
SimSceneTVB Perception is a corpus of 100 sound scenes of 45s each representing urban sound environments, including: 6 scenes recorded in Paris, 19 scenes simulated using simScene to replicate recorded scenarios, 75 scenes simulated using…
0 papers · 0 benchmarks
Simulacra Aesthetic Captions is a dataset of over 238000 synthetic images generated with AI models such as CompVis latent GLIDE and Stable Diffusion from over forty thousand user submitted prompts.
0 papers · 0 benchmarks
The SolProp dataset is a valuable resource for predicting the solubility limits of organic solutes in various solvents and at different temperatures.
0 papers · 0 benchmarks
The Sound Events for Surveillance Applications (SESA) dataset files were obtained from Freesound.
0 papers · 0 benchmarks
SourceData-NLP (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
Introduction: The scientific publishing landscape is expanding rapidly, creating challenges for researchers to stay up-to-date with the evolution of the literature.
0 papers · 0 benchmarks
The SoyCultivarVein dataset is a publicly available dataset, which comprises 100 categories (cultivars) with 6 samples (leaf images) in each cultivar and thus has a total number of 100×6 = 600 images (Yu et al.
0 papers · 0 benchmarks
Sparse LiDAR extracted from velodyne 64 beams in KITTI dataset.
0 papers · 0 benchmarks
A novel 360◦ fisheye panoramas dataset, i.e., the Spherical-Navi image dataset is collected, with a unique labeling strategy enabling automatic generation of an arbitrary number of negative samples (wrong heading direction).
0 papers · 0 benchmarks
SportsSum is a Chinese sports game summarization dataset that contains 5,428 soccer games of live commentaries and the corresponding news articles.
0 papers · 0 benchmarks
The dataset contains 256x256 tiles extracted from Whole Slide Images (WSI) of mouse liver tissue stained with H&E and Masson's Trichrome.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 3000+ originally Stair images captured and crowdsourced from over 500+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals at…
0 papers · 0 benchmarks
State Farm (State Farm Distracted Driver Detection)
该数据集是一个全面而多样化的驾驶员行为监测数据集,其中包括来自美洲,亚洲和非洲的26名不同种族,肤色和性别(13名男性和13名女性)的参与者。数据集中的所有图像都是由固定在汽车仪表板上的摄像头拍摄的,所有图像都是RGB像素。该数据集共包含22424张图像。
0 papers · 0 benchmarks
Student-Teacher Prompting is an instructional strategy used to guide a learner's behavior.
0 papers · 0 benchmarks
Sugar Beets 2016 is a robot dataset for plant classification as well as localization and mapping that covers the relevant stages for robotic intervention and weed control.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 7000+ original Suitcase/Luggage images captured and crowdsourced from over 800+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals…
0 papers · 0 benchmarks
SuperBench is a comprehensive evaluation system for large language models that includes five benchmark datasets: ExtremeGLUE for semantics, CodeBench for code, AlignBench for alignment, AgentBench for intelligent agents, and SafetyBench…
0 papers · 0 benchmarks
The SuperLim dataset is a Swedish version of the English benchmarking platform (Super)GLUE.
0 papers · 0 benchmarks
Supplementary material (Funding Covid-19 research: Insights from an exploratory analysis using open data infrastructures - Supplementary material)
Funding Covid-19 research: Insights from an exploratory analysis using open data infrastructures - Supplementary material
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.