Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 79 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3745–3792 of 12,172
A database of images of approximately 960 unique plants belonging to 12 species at several growth stages is made publicly available.
7 papers · 0 benchmarks
ataset format Each row in the dataset splits represents one instance and contains the following tab-separated columns: articleid - article id corresponding to the id of the claim in the LIAR dataset statement - the text of the claim author…
7 papers · 0 benchmarks
These are the protein-ligand complexes of the PoseBusters Benchmark set as described in the paper "PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences" [1] with associated code at…
7 papers · 0 benchmarks
PrideMM comprising 5,063 text-embedded images associated with the LGBTQ+ Pride movement
7 papers · 1 benchmark
RCooper (Roadside Cooperative Perception Dataset)
The first real-world, large-scale Roadside Cooperative Perception Dataset, RCooper, is released to bloom research on roadside cooperative perception for practical applications.
7 papers · 0 benchmarks
RL Unplugged is suite of benchmarks for offline reinforcement learning.
7 papers · 0 benchmarks
RWSD (The Winograd Schema Challenge (Russian))
A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its resolution.
7 papers · 1 benchmark
The dataset specifically focuses on the value of synthetic data to aid computer vision algorithms in their ability to automatically detect aircraft and their attributes in satellite imagery.
7 papers · 0 benchmarks
Reddit Corpus is part of a repository of conversational datasets consisting of hundreds of millions of examples, and a standardised evaluation procedure for conversational response selection models using '1-of-100 accuracy'.
7 papers · 0 benchmarks
RefMatte is the first large-scale challenging dataset under the task referring image matting, generated by a comprehensive image composition and expression generation engine on top of current public high-quality matting foregrounds with…
7 papers · 3 benchmarks
A data set containing citations, citation contexts, and papers.
7 papers · 0 benchmarks
Understanding spatial relations (e.g., “laptop on table”) in visual input is important for both humans and robots.
7 papers · 1 benchmark
A dataset of over 65,000 pairs of incorrectly white-balanced images and their corresponding correctly white-balanced images.
7 papers · 0 benchmarks
The Retrieval-SFM dataset is used for instance image retrieval.
7 papers · 0 benchmarks
RuBQ (Russian Knowledge Base Questions)
The first Russian knowledge base question answering (KBQA) dataset.
7 papers · 0 benchmarks
Synthetic COCO (S-COCO) is a synthetically created dataset for homography estimation learning.
7 papers · 1 benchmark
SEN12MS-CR-TS is a multi-modal and multi-temporal data set for cloud removal.
7 papers · 1 benchmark
Test set version 1 for the San Francisco eXtra Large dataset
7 papers · 1 benchmark
SI-SCORE (Synthetic Interventions on Scenes for Robustness Evaluation)
A synthetic dataset uses for a systematic analysis across common factors of variation.
7 papers · 0 benchmarks
English subset of the SLAKE dataset, comprising 642 images and more than 7,000 question–answer pairs.
7 papers · 0 benchmarks
SME (Standard Multimodal Explanation)
SME is a new dataset for Multi-modal Explanation for Visual Question Answering comprising 1,028,230 samples, with 1,656 visual objects requiring detection in explanations.
7 papers · 1 benchmark
SNARE, short for ShapeNet Annotated with Referring Expressions, is a benchmark requires a model to choose which of two objects is being referenced by a natural language description.
7 papers · 0 benchmarks
SODA-A is a large-scale benchmark specialized for small object detection task under aerial scenes, which has 800203 instances with oriented rectangle box annotation across 9 classes.
7 papers · 0 benchmarks
SSC (Spiking Speech Commands v0.2)
The SSC dataset is a spiking version of the Speech Commands dataset release by Google (Speech Commands).
7 papers · 1 benchmark
Cross-view Image Dataset Across Drone and Satellite - multi-height - multi-scene
7 papers · 0 benchmarks
SVD (Short Video Dataset)
SVD is a large-scale short video dataset, which contains over 500,000 short videos collected from http://www.douyin.com and over 30,000 labeled pairs of near-duplicate videos.
7 papers · 0 benchmarks
Evaluating human-scene interaction requires precise annotations for camera pose and scene geometry.
7 papers · 1 benchmark
SingFake (SingFake: Singing Voice Deepfake Detection)
The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage.
7 papers · 0 benchmarks
A multimodal agent benchmark on professional data science and engineering.
7 papers · 0 benchmarks
The Stanford Light Field Archive is a collection of several light fields for research in computer graphics and vision.
7 papers · 0 benchmarks
To comprehensively evaluate the effectiveness and generalization ability of style transfer methods, we build StyleBench that covers 73 distinct styles, ranging from paintings, flat illustrations, 3D rendering to sculptures with varying…
7 papers · 1 benchmark
We construct a style-balanced dataset, called StyleGallery, covering several open source datasets.
7 papers · 0 benchmarks
The Sunnybrook Cardiac Data (SCD), also known as the 2009 Cardiac MR Left Ventricle Segmentation Challenge data, consist of 45 cine-MRI images from a mixed of patients and pathologies: healthy, hypertrophy, heart failure with infarction…
7 papers · 0 benchmarks
SynPick is a synthetic dataset for dynamic scene understanding in bin-picking scenarios.
7 papers · 1 benchmark
The SynthHands dataset is a dataset for hand pose estimation which consists of real captured hand motion retargeted to a virtual hand with natural backgrounds and interactions with different objects.
7 papers · 0 benchmarks
The TACRED-Revisited dataset improves the crowd-sourced TACRED dataset for relation extraction by relabeling the dev and test sets using expert linguistic annotators.
7 papers · 1 benchmark
TAPE (resToration of digitized Analog videotaPEs)
A dataset of videos synthetically degraded with Adobe After Effects to exhibit artifacts resembling those of real-world analog videotapes.
7 papers · 1 benchmark
TASD (Target Aspect Sentiment Detection)
Aspect-based sentiment analysis (ABSA) aims to detect the targets (which are composed by continuous words), aspects and sentiment polarities in text.
7 papers · 1 benchmark
TERRa (Textual Entailment Recognition for Russian)
Textual Entailment Recognition has been proposed recently as a generic task that captures major semantic inference needs across many NLP applications, such as Question Answering, Information Retrieval, Information Extraction, and Text…
7 papers · 1 benchmark
Recent work about synthetic indoor datasets from perspective views has shown significant improvements of object detection results with Convolutional Neural Networks(CNNs).
7 papers · 0 benchmarks
TIAGE is a topic-shift aware dialog benchmark constructed utilizing human annotations on topic shifts.
7 papers · 0 benchmarks
TREK-150 is a benchmark dataset for object tracking in First Person Vision (FPV) videos composed of 150 densely annotated video sequences.
7 papers · 0 benchmarks
The TUM Kitchen dataset is an action recognition dataset that contains 20 video sequences captured by 4 cameras with overlapping views.
7 papers · 0 benchmarks
A sentiment analysis Tunisian Arabizi Dataset, collected from social networks, preprocessed for analytical studies and annotated manually by Tunisian native speakers.
7 papers · 0 benchmarks
TUT-SED Synthetic 2016 contains of mixture signals artificially generated from isolated sound events samples.
7 papers · 0 benchmarks
This dataset contains simulations of a complex, large-scale chemical plant proposed by Downs and Vogel (1993).
7 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.