Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 69 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3265–3312 of 12,172
This is a multiscale dynamic human mobility flow dataset across the United States, with data starting from January 1st, 2019.
9 papers · 0 benchmarks
MuSe-CaR (Multimodal Sentiment Analysis in Car Reviews)
The MuSe-CAR database is a large, multimodal (video, audio, and text) dataset which has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that…
9 papers · 0 benchmarks
NCLS (Neural Cross-Lingual Summarization Corpora)
Presents two high-quality large-scale CLS datasets based on existing monolingual summarization datasets.
9 papers · 0 benchmarks
NELA-GT-2018 is a dataset for the study of misinformation that consists of 713k articles collected between 02/2018-11/2018.
9 papers · 0 benchmarks
NICO (Non-I.I.D. Image dataset with Contexts)
I.I.D.
9 papers · 2 benchmarks
A popular dataset for node classification on heterogeneous graphs.
9 papers · 1 benchmark
OASIS-1 (Open Access Series of Imaging Studies)
The Open Access Series of Imaging Studies (OASIS) is a project aimed at making neuroimaging data sets of the brain freely available to the scientific community.
9 papers · 0 benchmarks
Presents half a million samples and structured meta-data to encourage further research and societal engagement.
9 papers · 1 benchmark
OntoGUM is an OntoNotes-like coreference dataset converted from GUM, an English corpus covering 12 genres using deterministic rules.
9 papers · 1 benchmark
PEC (Persona-Based Empathetic Conversational)
A novel large-scale multi-domain dataset for persona-based empathetic conversations.
9 papers · 0 benchmarks
PPM is a portrait matting benchmark with the following characteristics: - Fine Annotation - All images are labeled and checked carefully.
9 papers · 1 benchmark
PSI-AVA is a dataset designed for holistic surgical scene understanding.
9 papers · 0 benchmarks
PTB-TIR is a Thermal InfraRed (TIR) pedestrian tracking benchmark, which provides 60 TIR sequences with mannuly annoations.
9 papers · 0 benchmarks
The Pascal Panoptic Parts dataset consists of annotations for the part-aware panoptic segmentation task on the PASCAL VOC 2010 dataset.
9 papers · 2 benchmarks
PathTrack is a dataset for person tracking which contains more than 15,000 person trajectories in 720 sequences.
9 papers · 0 benchmarks
PerSeg is a dataset for personalized segmentation.
9 papers · 1 benchmark
PhenoBench (PhenoBench — A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain)
The PhenoBench dataset contains multiple image segmentation challenges from the agricultural domain.
9 papers · 0 benchmarks
Most existing MOT datasets are captured using pinhole cameras, which are characterized by a narrow-FoV and linear sensor motion.
9 papers · 1 benchmark
Consists of 330,000 sketches and 204,000 photos spanning across 110 categories.
9 papers · 0 benchmarks
RADDet (Range-Azimuth-Doppler based Radar Dataset)
RADDet is a radar dataset that contains radar data in the form of Range-Azimuth-Doppler tensors along with the bounding boxes on the tensor for dynamic road users, category labels, and 2D bounding boxes on the Cartesian Bird-Eye-View range…
9 papers · 0 benchmarks
RAVEN-FAIR is a modified version of the RAVEN dataset.
9 papers · 0 benchmarks
The RU-APC (Rutgers APC) dataset is a valuable resource for researchers and developers working on robotic perception solutions for warehouse picking challenges.
9 papers · 0 benchmarks
RadQA (A Question Answering Dataset to Improve Comprehension of Radiology Reports)
RadQA is a radiology question answering dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans (resulting in a total of 6148 question-answer evidence pairs) by physicians.
9 papers · 1 benchmark
A human-curated ChineseReading Comprehension dataset on Opinion.
9 papers · 0 benchmarks
ReDWeb-S is a large-scale challenging dataset for Salient Object Detection.
9 papers · 0 benchmarks
Roadside Perception 3D (Rope3D) is a dataset for autonomous driving and monocular 3D object detection task consisting of 50k images and over 1.5M 3D objects in various scenes, which are captured under different settings including various…
9 papers · 1 benchmark
The SALMon dataset and benchmark was introduced in the paper "A Suite for Acoustic Language Model Evaluation", with the goal of evaluating the modelling abilities of speech language models with regards to different kinds of acoustic…
9 papers · 1 benchmark
SAMRS is a remote sensing segmentation dataset which provides object category, location, and instance information that can be used for semantic segmentation, instance segmentation, and object detection, either individually or in…
9 papers · 0 benchmarks
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.
9 papers · 0 benchmarks
SHERLOCK is a corpus of 363K commonsense inferences grounded in 103K images.
9 papers · 0 benchmarks
SOBA (Shadow-OBject Association)
A new dataset called SOBA, named after Shadow-OBject Association, with 3,623 pairs of shadow and object instances in 1,000 photos, each with individual labeled masks.
9 papers · 1 benchmark
SPARTQA (SPAtial Reasoning on Textual Question Answering)
SpartQA is a textual question answering benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior datasets and that is challenging for state-of-the-art language models…
9 papers · 0 benchmarks
SPARTQA - (SPAtial Reasoning on Textual Question Answering.)
We take advantage of the ground truth of NLVR images, design CFGs to generate stories, and use spatial reasoning rules to ask and answer spatial reasoning questions.
9 papers · 0 benchmarks
SPEECH-COCO contains speech captions that are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.
9 papers · 0 benchmarks
Provides four new test sets for the Stanford Question Answering Dataset (SQuAD) and evaluate the ability of question-answering systems to generalize to new data.
9 papers · 3 benchmarks
SUM is a new benchmark dataset of semantic urban meshes which covers about 4 km2 in Helsinki (Finland), with six classes: Ground, Vegetation, Building, Water, Vehicle, and Boat.
9 papers · 0 benchmarks
SciRepEval is a comprehensive benchmark for training and evaluating scientific document representations.
9 papers · 0 benchmarks
The SciTail dataset is an entailment dataset created from multiple-choice science exams and web sentences.
9 papers · 1 benchmark
SelQA is a dataset that consists of questions generated through crowdsourcing and sentence length answers that are drawn from the ten most prevalent topics in the English Wikipedia.
9 papers · 0 benchmarks
A large-scale dataset for the point cloud completion task on the ShapeNet dataset.
9 papers · 1 benchmark
SketchyScene is a large-scale dataset of scene sketches to advance research on sketch understanding at both the object and scene level.
9 papers · 0 benchmarks
The Standardized Project Gutenberg Corpus (SPGC) is an open science approach to a curated version of the complete PG data containing more than 50,000 books and more than 3×109 word-tokens.
9 papers · 0 benchmarks
SwissDial is an annotated parallel corpus of spoken Swiss German across 8 major dialects, plus a Standard German reference.
9 papers · 0 benchmarks
TEMPO (Localizing Moments in Video with Temporal Language)
TEMPOral reasoning in video and language (TEMPO) is a dataset that consists of two parts: a dataset with real videos and template sentences (TEMPO - Template Language) which allows for controlled studies on temporal language, and a human…
9 papers · 0 benchmarks
TRIP (Tiered Reasoning for Intuitive Physics)
Tiered Reasoning for Intuitive Physics (TRIP) is a novel commonsense reasoning dataset with dense annotations that enable multi-tiered evaluation of machines’ reasoning process.
9 papers · 0 benchmarks
TSAC (Tunisian Sentiment Analysis Corpus)
Tunisian Sentiment Analysis Corpus (TSAC) is a Tunisian Dialect corpus of 17.000 comments from Facebook.
9 papers · 0 benchmarks
TSSB (Time Series Segmentation Benchmark)
The time series segmentation benchmark (TSSB) currently contains 75 annotated time series (TS) with 1-9 segments.
9 papers · 1 benchmark
TUM-GAID (TUM Gait from Audio, Image and Depth) collects 305 subjects performing two walking trajectories in an indoor environment.
9 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.