Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 160 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7633–7680 of 12,172
The dataset contains 73 satellite images of different forests damaged by wildfires across Europe with a resolution of up to 10m per pixel.
1 paper · 1 benchmark
Original images and images with RUSTICO filters applied Also a csv with classes is included
1 paper · 1 benchmark
Transit agencies use the General Transit Feed Specification (GTFS) to publish transit data.
1 paper · 0 benchmarks
This dataset contains the bus trajectory dataset collected by 6 volunteers who were asked to travel across the sub-urban city of Durgapur, India, on intra-city buses (route name: 54 Feet).
1 paper · 0 benchmarks
Business license datasets and source code for named entity recognition.
1 paper · 0 benchmarks
This is a proprietary dataset from a large internet services company of ranked pairs of relevant and irrelevant businesses for different queries, for a total of 17,069 pairs.
1 paper · 0 benchmarks
We scraped the 53 most popular C# repositories from GitHub and extracted all commits since the beginning of the project’s history.
1 paper · 1 benchmark
This is the C++ dataset used in the TASTY research paper which was published at the ICLR DL4Code (Deep Learning for Code) workshop.
1 paper · 0 benchmarks
The feature files are named with the youtube IDs.
1 paper · 0 benchmarks
📚 CADBench CADBench is a comprehensive benchmark to evaluate the ability of LLMs to generate CAD scripts.
1 paper · 0 benchmarks
CAD-EdgeTune dataset is acquired using a Husarion ROSbot 2.0 and ROSbot 2.0 Pro with the collection speed set to 5 frames per second from a suburban university environment.
1 paper · 0 benchmarks
We introduce the CADNET dataset, which is an annotated collection of 3,317 3D Engineering models over 43 categories.
1 paper · 0 benchmarks
13,201 clips from 79 TV shows.
1 paper · 2 benchmarks
This dataset labeled by SAR experts was created using 102 Chinese Gaofen-3 images and 108 Sentinel-1 images.
1 paper · 0 benchmarks
CAFD (Central Asian Food Dataset)
This dataset comprises 16 499 images with 42 classes encompassing the most popular Central Asian cuisine consumed locally.
1 paper · 0 benchmarks
CAGUI (Chinese Android GUI Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CAL10K (Computer Audition Lab 10000)
The CAL10K dataset (introduced as Swat10k) contains 10,870 songs that are weakly-labelled using a tag vocabulary of 475 acoustic tags and 153 genre tags.
1 paper · 0 benchmarks
The CAL500 Expansion (CAL500exp) dataset is an enriched version of the CAL500 music information retrieval dataset.
1 paper · 0 benchmarks
CAMO-FS Dataset comes with the paper entitled The Art of Camouflage: Few-shot Learning for Animal Detection and Segmentation.
1 paper · 2 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
CAP-DATA is a large-scale benchmark consisting of 11,727 in-the-wild accident videos with over 2.19 million frames together with labeled fact-effect-reason-introspection description and temporal accident frame label.
1 paper · 0 benchmarks
https://drive.google.com/file/d/1XJTfD8Ch-IxmG5VHtKkxGZT336Fl1Q/view?usp=drivelink
1 paper · 0 benchmarks
This dataset contains synthetic images extracted from the CARLA simulator along with rich information extracted from the deferred rendering pipeline of Unreal Engine 4.
1 paper · 0 benchmarks
CARLE (Cellular Automata Reinforcement Learning Environment)
CARLE is a life-like cellular automata simulator and reinforcement learning environment.
1 paper · 0 benchmarks
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data.
1 paper · 3 benchmarks
CASA (Clinical annotations for automatic Stuttering Assessment)
Annotation Guidlines Three speech and language pathologists, with experience ranging from 2 to 40 years, independently annotated and analyzed the audiovisual samples sourced from the Fluencybank Adults Who Stutter(AWS) dataset.
1 paper · 0 benchmarks
CASCONet is a a collection of data about the CAS Conference (CASCON) for the past 25 years including information about papers, technology showcase demos, workshops, and keynote presentations.
1 paper · 0 benchmarks
Annotation corpus of cybersecurity event in news articles The corpus contains 1000 annotation and source files.
1 paper · 0 benchmarks
CASP13 MQA is a dataset that contains predicted models for CASP13 targets and their scores.
1 paper · 0 benchmarks
CASR (Cyclist Arm Signal Recognition)
CASR is a dataset for cyclist arm signal recognition in videos.
1 paper · 0 benchmarks
The CASTLE Benchmark is a comprehensive dataset and a scoring method for evaluating single or combinations of static analyzers with a focus on security.
1 paper · 0 benchmarks
CAShift (Cloud Attack & Normality Shift Dataset)
CAShift is the first multiple normality shift-aware Log-Based Anomaly Detection (LAD) dataset specifically designed for cloud systems, which considers different software roles in cloud systems and attack behavior among cloud components.
1 paper · 0 benchmarks
CAT (Context Adjustment Training)
CAT is a specialized dataset for co-saliency detection - one of the core tasks in the field of computer vision.
1 paper · 0 benchmarks
CAT is a specialized dataset for co-saliency detection.
1 paper · 0 benchmarks
The CATH (Class, Architecture, Topology, Homology) [65] database is a comprehensive resource for protein structure classification that hierarchical group proteins based on their structural features.
1 paper · 1 benchmark
The CATH (Class, Architecture, Topology, Homology) [65] database is a comprehensive resource for protein structure classification that hierarchical group proteins based on their structural features.
1 paper · 1 benchmark
CAsT-answerability dataset contains binary answerability labels on three levels: sentence, passage, and ranking.
1 paper · 0 benchmarks
CBTex (Synthetic CardBoard Textures)
Dataset of >200 synthetic cardboard texture images that were rendered with DoubeGum's cardboard shader in Blender.
1 paper · 0 benchmarks
A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenario
1 paper · 1 benchmark
CC-Riddle is a Chinese character riddle dataset covering the majority of common simplified Chinese characters by crawling riddles from the Web and generating brand new ones.
1 paper · 0 benchmarks
CCEDD (cervical cell edge detection dataset)
We construct a largest publicly Cervical Cell Edge Detection Dataset (CCEDD) based on our Local Label Point Correction (LLPC).
1 paper · 0 benchmarks
CCGR (Cross-Covariate Gait Recognition)
CCGR (Cross-Covariate Gait Recognition), the first gait dataset for studying cross-covariate challenges, contains 970 subjects, approximately 1.6 million sequences, 53 covariates, and 33 views.
1 paper · 0 benchmarks
To address the scarcity of high-quality safety datasets in the Chinese, we open-sourced the CCI (Chinese Corpora Internet) dataset on November 29, 2023.
1 paper · 0 benchmarks
CCIHP (Characterized Crowd Instance-level Human Parsing)
CCIHP dataset is devoted to fine-grained description of people in the wild with localized & characterized semantic attributes.
1 paper · 0 benchmarks
CCPT (Conceptual Combination with Property Type)
CCPT is a dataset containing 12.3K triplets of noun phrases, properties, and property types for conceptual combination.
1 paper · 0 benchmarks
CCQA is new web-scale dataset for in-domain model pre-training.
1 paper · 0 benchmarks
CCSE (Chinese Character Stroke Extraction)
Chinese Character Stroke Extraction (CCSE) is a benchmark containing two large-scale datasets: Kaiti CCSE (CCSE-Kai) and Handwritten CCSE (CCSE-HW).
1 paper · 0 benchmarks
The CSUST Chinese Traffic Sign Detection Benchmark (CCTSDB) is an existing dataset for traffic sign detection.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.