Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 94 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4465–4512 of 12,172
The semantic line (SEL) dataset contains 1,750 outdoor images in total, which are split into 1,575 training and 175 testing images.
5 papers · 1 benchmark
SERV-CT (SERV-CT: A disparity dataset from CT for validation of endoscopic 3D reconstruction)
Endoscopic stereo reconstruction for surgical scenes gives rise to specific problems, including the lack of clear corner features, highly specular surface properties, and the presence of blood and smoke.
5 papers · 0 benchmarks
a high-level explanation of the dataset characteristics explain motivations and summary of its content potential use cases of the dataset
5 papers · 1 benchmark
SODA-D is a large-scale dataset tailored for small object detection in driving scenario, which is built on top of MVD dataset and owned data, where the former is a dataset dedicated to pixel-level understanding of street scenes, and the…
5 papers · 1 benchmark
The SOFC-Exp corpus contains 45 scientific publications about solid oxide fuel cells (SOFCs), published between 2013 and 2019 as open-access articles all with a CC-BY license.
5 papers · 0 benchmarks
A large-scale evaluation set that provides human ratings for the plausibility of 10,000 SP pairs over five SP relations, covering 2,500 most frequent verbs, nouns, and adjectives in American English.
5 papers · 0 benchmarks
The dataset for the SPHERE challenge consists on a multimodal activity recognition dataset consisting of accelerometer, RGB-D and environmental data.
5 papers · 0 benchmarks
SSN (Semantic Scholar Network)
SSN (short for Semantic Scholar Network) is a scientific papers summarization dataset which contains 141K research papers in different domains and 661K citation relationships.
5 papers · 0 benchmarks
STACKEX expands beyond the only existing genre (i.e., academic writing) in keyphrase generation tasks.
5 papers · 0 benchmarks
SWDE (Structured Web Data Extraction)
This dataset is a real-world web page collection used for research on the automatic extraction of structured data (e.g., attribute-value pairs of entities) from the Web.
5 papers · 1 benchmark
SWINSEG (Singapore Whole sky Nighttime Image SEGmentation Database)
The SWINSEG dataset contains 115 nighttime images of sky/cloud patches along with their corresponding binary ground truth maps.
5 papers · 1 benchmark
SYNS-Patches dataset, which is a subset of SYNS.
5 papers · 0 benchmarks
Specially designed to evaluate active learning for video object detection in road scenes.
5 papers · 0 benchmarks
Sales (Rossmann Store Sales)
Forecast Sales using ARIMA and SARIMA
5 papers · 0 benchmarks
The Sarcasm Corpus contains sarcastic and non-sarcastic utterances of three different types, which are balanced with half of the samples being sarcastic and half non-sarcastic.
5 papers · 0 benchmarks
Sewer-ML is a sewer defect dataset.
5 papers · 0 benchmarks
X-ray images in this data set have been collected by Shenzhen No.3 Hospital in Shenzhen, Guangdong providence,China.
5 papers · 0 benchmarks
Simitate is a hybrid benchmarking suite targeting the evaluation of approaches for imitation learning.
5 papers · 0 benchmarks
The Sims4Action Dataset: a videogame-based dataset for Synthetic→Real domain adaptation for human activity recognition.
5 papers · 0 benchmarks
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
5 papers · 0 benchmarks
SoMeSci (Software Mentions in Scientific Articles)
Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling.
5 papers · 0 benchmarks
Comprises of 171,191 video segments from 346 high-quality soccer games.
5 papers · 0 benchmarks
The Song Describer Dataset (SDD) contains ~1.1k captions for 706 permissively licensed music recordings.
5 papers · 1 benchmark
SpaRTUN a dataset synthesized for transfer learning on spatial question answering (SQA) and spatial role labeling (SpRL).
5 papers · 0 benchmarks
SpaceNet 1: Building Detection v1 is a dataset for building footprint detection.
5 papers · 2 benchmarks
A multilingual image dataset with spatial relation annotations and object features for image-to-text generation, built using 2,026 images from the PASCAL VOC2008 dataset.
5 papers · 0 benchmarks
Species-800 is a corpus for species entities, which is based on manually annotated abstracts.
5 papers · 1 benchmark
- Games dataset containing 100,000 Gameplay Images of 175 Video Games across 10 Sports Genres - AMERICAN FOOTBALL, BASKETBALL, BIKE RACING, CAR RACING, FIGHTING, HOCKEY, SOCCER, TABLE TENNIS, TENNIS.
5 papers · 2 benchmarks
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks.
5 papers · 0 benchmarks
StreetStyle is a large-scale dataset of photos of people annotated with clothing attributes, and use this dataset to train attribute classifiers via deep learning.
5 papers · 0 benchmarks
The SUGARCREPE++ dataset evaluates the sensitivity of vision language models (VLMs) and unimodal language models (ULMs) to semantic and lexical alterations.
5 papers · 0 benchmarks
SummEdits is a benchmark designed to measure the ability of Large Language Models (LLMs) to reason about facts and detect inconsistencies.
5 papers · 0 benchmarks
In the Learning to Summarize from Human Feedback paper, a reward model was trained from human feedback.
5 papers · 0 benchmarks
SyRIP is a hybrid synthetic and real infant pose (SyRIP) dataset with small yet diverse real infant images as well as generated synthetic infant poses and (2) a multi-stage invariant representation learning strategy that could transfer the…
5 papers · 0 benchmarks
T2Ranking is a large-scale Chinese benchmark for passage ranking.
5 papers · 0 benchmarks
TAS500 is a semantic segmentation dataset for autonomous driving in unstructured environments.
5 papers · 0 benchmarks
This data collection consists of images acquired during chemoradiotherapy of 20 locally-advanced, non-small cell lung cancer patients.
5 papers · 0 benchmarks
Inolves an annotated a large number of cells, including normal epithelial and myoepithelial breast cells (localized in ducts and lobules), invasive carcinomatous cells, fibroblasts, endothelial cells, adipocytes, macrophages and…
5 papers · 1 benchmark
TRANCE (Transformation Driven Visual Reasoning)
TRANCE extends CLEVR by asking a uniform question, i.e.
5 papers · 0 benchmarks
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
This is a benchmark set for Traveling salesman problem (TSP) with characteristics that are different from the existing benchmark sets.
5 papers · 1 benchmark
Taiga is a corpus, where text sources and their meta-information are collected according to popular ML tasks.
5 papers · 0 benchmarks
This is a 16.2-million frame (50-hour) multimodal dataset of two-person face-to-face spontaneous conversations.
5 papers · 0 benchmarks
The Taskmaster-2 dataset consists of 17,289 dialogs in seven domains: restaurants (3276), food ordering (1050), movies (3047), hotels (2355), flights (2481), music (1602), and sports (3478).
5 papers · 0 benchmarks
A collection of 2511 recipes for zero-shot learning, recognition and anticipation.
5 papers · 0 benchmarks
Tencent ML-Images is a large open-source multi-label image database, including 17,609,752 training and 88,739 validation image URLs, which are annotated with up to 11,166 categories.
5 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.