Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 144 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6865–6912 of 12,172
SQL-Eval is an open-source PostgreSQL evaluation dataset released by Defog, constructed based on Spider.
2 papers · 1 benchmark
SR-Reg (SynthRAD Registration)
SR-Reg is a brain MR-CT registration dataset, deriving from SynthRAD 2023 (https://synthrad2023.grand-challenge.org/).
2 papers · 1 benchmark
SSD (Sub-Slot Dialogue dataset)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
2 papers · 0 benchmarks
SSD_PHONE (Sub-Slot Dialogue dataset phone domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
2 papers · 0 benchmarks
SSL4EO-S12 is a large-scale, global, multimodal, and multi-seasonal corpus of satellite imagery from the ESA Sentinel-1 & -2 satellite missions.
2 papers · 0 benchmarks
STIR (Scaled and Translated Image Recognition)
While convolutions are known to be invariant to (discrete) translations, scaling continues to be a challenge and most image recognition networks are not invariant to them.
2 papers · 0 benchmarks
SUMS (Summit Vitals: Multi-Camera and Multi-Signal Biosensing at High Altitudes)
Here is SUMS dataset collected by Qinghai University.
2 papers · 0 benchmarks
SVBench (Streaming Video Understanding Benchmark)
Dataset Card for SVBench This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
SWAX (Sense Wax Attack dataset)
Comprised of real human and wax figure images and videos that endorse the problem of face spoofing detection.
2 papers · 0 benchmarks
Saint Gall dataset contains handwritten historical manuscripts written in Latin that date back to the 9th century.
2 papers · 1 benchmark
Sakuga-42M is a large-scale hand-drawn cartoon video dataset for academic research purposes, it comprises 42 million cartoon keyframes covering various artistic styles, regions, and years, with comprehensive semantic annotations including…
2 papers · 0 benchmarks
A salient object subitizing image dataset of about 14K everyday images which are annotated using an online crowdsourcing marketplace.
2 papers · 0 benchmarks
Sangraha is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large-scale translations.
2 papers · 0 benchmarks
The resources for this dataset can be found at https://www.openml.org/d/182 Author: Ashwin Srinivasan, Department of Statistics and Data Modeling, University of Strathclyde Source: UCI - 1993 Please cite: UCI The database consists of the…
2 papers · 0 benchmarks
A dataset of ranked scan-CAD similarity annotations, enabling new, fine-grained evaluation of CAD model retrieval to cluttered, noisy, partial scans.
2 papers · 0 benchmarks
ScanBank is a benchmark dataset for figure extraction from scanned electronic theses and dissertations containing 10 thousand scanned page images, manually labeled by humans as to the presence of the 3.3 thousand figures or tables found…
2 papers · 0 benchmarks
This is a dataset of paired OpenAlex authorids (https://docs.openalex.org/about-the-data/author) and tweeterid.
2 papers · 0 benchmarks
SciGen is a challenge dataset for the task of reasoning-aware data-to-text generation consisting of tables from scientific articles and their corresponding descriptions.
2 papers · 0 benchmarks
SciHTC is a dataset for hierarchical multi-label text classification (HMLTC) of scientific papers which contains 186,160 papers and 1,233 categories from the ACM CCS tree.
2 papers · 0 benchmarks
SegPANDA200 (Segmentation task on PANDA challenge in 200 microns by 512px)
SegPANDA200 is a public pathological H&E image dataset from segmentation task on PANDA challenge in 200 microns by 512px made in the same manner from PANDA challenge dataset .
2 papers · 0 benchmarks
Pre-training is a strong strategy for enhancing visual models to efficiently train them with a limited number of labeled images.
2 papers · 0 benchmarks
The SegmentedTables dataset is a collection of almost 2,000 tables extracted from 352 machine learning papers.
2 papers · 0 benchmarks
SemClinBr (A multi‑institutional and multi‑specialty semantically annotated corpus for Portuguese clinical NLP tasks)
Background: The high volume of research focusing on extracting patient information from electronic health records (EHRs) has led to an increase in the demand for annotated corpora, which are a precious resource for both the development and…
2 papers · 1 benchmark
NSURL-2019 Shared Task 8: Semantic Question Similarity in Arabic This dataset contains 11,997 pairs of questions in MSA Arabic that are assigned either a label of 0, for no semantic similarity, or 1 otherwise.
2 papers · 0 benchmarks
Semantic Trails Datasets (STDs) are two different datasets of semantically annotated trails created starting from check-ins performed on the Foursquare social network.
2 papers · 0 benchmarks
Homepage | GitHub LiDARs are one of the main sensors used for autonomous driving applications, providing accurate depth estimation regardless of lighting conditions.
2 papers · 0 benchmarks
The Sentimental LIAR dataset is a modified and further extended version of the LIAR extension introduced by Kirilin et al.
2 papers · 0 benchmarks
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Large-scale shadows from buildings in a city play an important role in determining the environmental quality of public spaces.
2 papers · 0 benchmarks
The synthetic ShapeNet intrinsic image decomposition dataset used for training the deep CNN models IntrinsicNet and RetiNet of CVPR2018.
2 papers · 0 benchmarks
The SheetCopilot dataset contains 28 evaluation workbooks and 221 spreadsheet manipulation tasks that are applied to these workbooks.
2 papers · 1 benchmark
A newly developed natural scene text dataset of Chinese shop signs in street views.
2 papers · 0 benchmarks
Recent applications of LLMs in Machine Reading Comprehension (MRC) systems have shown impressive results, but the use of shortcuts, mechanisms triggered by features spuriously correlated to the true label, has emerged as a potential threat…
2 papers · 0 benchmarks
A dataset of 53 complex-valued signal modulation classes.
2 papers · 0 benchmarks
The Signal Media One-Million News Articles Dataset dataset by Signal Media was released to facilitate researching news articles.
2 papers · 0 benchmarks
A new benchmark dataset for simple question answering over knowledge graphs that was created by mapping SimpleQuestions entities and predicates from Freebase to DBpedia.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
LLMs' lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarcity of relevant data.
2 papers · 0 benchmarks
A curated and 3-D pose-annotated subset of RGB videos sourced from Kinetics-700, a large-scale action dataset.
2 papers · 1 benchmark
A dataset derived from the recently introduced Mimetics dataset.
2 papers · 2 benchmarks
Collects a huge number of job descriptions from Dice.com - one of the most popular career website about Tech jobs in USA.
2 papers · 0 benchmarks
SkyCam dataset is a collection of sky images from a variety of locations with diverse topological characteristics (Swiss Jura, Plateau and Pre-Alps regions), from both single and stereo camera settings coupled with a high-accuracy…
2 papers · 0 benchmarks
SmartCity consists of 50 images in total collected from ten city scenes including office entrance, sidewalk, atrium, shopping mall etc..
2 papers · 0 benchmarks
Online social platforms serve a critical role for individuals as they seek to fill informational and emotional needs, from informational support like advice to emotional support like expressions of sympathy, frequently by interacting with…
2 papers · 0 benchmarks
Data set constructed from YouTube comments (72,098 comments posted by 43,859 users on 623 relevant videos to the crisis)
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.