Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 163 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7777–7824 of 12,172
CROSS (Cross-Reference Omnidirectional Stitching IQA)
Cross-Reference Omnidirectional Stitching IQA is a novel omnidirectional image dataset containing stitched images as well as dual-fisheye images captured from standard quarters of 0◦, 90◦ , 180◦ and 270◦.
1 paper · 0 benchmarks
CRSB (Context Retrieval Supervision Benchmark)
The Official dataset proposed int the paper Context Awareness Gate For Retrieval Augmented Generation
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CSAbstruct is a new dataset of annotated computer science abstracts with sentence labels according to their rhetorical roles.
1 paper · 0 benchmarks
CSI is a criminal conversational dataset for speaker identification built from the CSI television show.
1 paper · 0 benchmarks
A daily emerging stock market dataset (Chinese CSI 300 dataset) including 300 stocks and 5,088 time steps from the CSMAR database.
1 paper · 1 benchmark
The raw .mat data collected in two different scenarios are provided.
1 paper · 0 benchmarks
The dataset contains gold-standard summary labels for 39 "CSI: Crime Scene Investigation" episodes from seasons 1-5.
1 paper · 0 benchmarks
We present CSL, a large-scale Chinese Scientific Literature dataset, which contains the titles, abstracts, keywords and academic fields of 396,209 papers.
1 paper · 0 benchmarks
A large-scale gloss-free sign language translation dataset with 1,985 hours of videos, approximately 86 times larger than the previous CSL-Daily dataset.
1 paper · 0 benchmarks
CSPRD (Chinese Stock Policy Retrieval Dataset)
The Chinese Stock Policy Retrieval Dataset (CSPRD) contains a Chinese policy corpus of 10,002 articles and 709 prospectus examples from 545 companies listed on China’s Science and Technology Innovation Board (STAR Market).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CSTS (Correlation Structures in Time Series)
CSTS: Correlation Structures in Time Series CSTS is a comprehensive synthetic benchmarking dataset designed specifically for evaluating correlation structure discovery in time series data.
1 paper · 0 benchmarks
Over 20,000 annotated synthetic images and web-scraped images of bicyclists with bounding box annotations in Pascal VOC format.
1 paper · 1 benchmark
CTFW is a large annotated procedural text dataset in the cybersecurity domain (3154 documents).
1 paper · 0 benchmarks
This dataset contains samples of CTI (Cyber Threat Intelligence) data in natural language, labeled with the corresponding adversarial techniques from the MITRE ATT&CK framework.
1 paper · 0 benchmarks
This dataset includes 720 directional B-format RIRs, i.e.
1 paper · 0 benchmarks
CUCO Database (A voice and speech corpus of patients who underwent upper airway surgery in pre-and post-operative states)
Many research articles have explored the impact of surgical interventions on voice and speech evaluations, but advances are limited by the lack of publicly accessible datasets.
1 paper · 0 benchmarks
The CUHK Face Alignment Database is dataset with 13,466 face images, among which 5, 590 images are from LFW and the remaining 7, 876 images are downloaded from the web.
1 paper · 0 benchmarks
CUHK-QA is a dataset for natural language-based person search using iterative questioning.
1 paper · 0 benchmarks
CUHK01 (CUHK Person Re-identification)
This dataset contains 971 identities from two disjoint camera views.
1 paper · 0 benchmarks
CURE (A dataset for Clinical Understanding & Retrieval Evaluation)
CURE is a retrieval dataset with a monolingual and two cross-lingual conditions, with splits spanning ten medical domains.
1 paper · 0 benchmarks
This is the dataset released along with the publication: CUTS: A Deep Learning and Topological Framework for Multigranular Unsupervised Medical Image Segmentation [[ArXiv]](https://arxiv.org/abs/2209.11359)…
1 paper · 0 benchmarks
CVE (Common Vulnerabilities and Exposures)
CVE stands for Common Vulnerabilities and Exposures.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CVIAN is a cross-view dataset to support geolocalization and disaster mapping with street-view and very high resolution (VHR) satellite imagery in Florida, USA after Hurricane IAN in 2022.
1 paper · 0 benchmarks
CVR (Congressional Voting Records Data Set)
This data set includes votes for each of the U.S.
1 paper · 1 benchmark
In this dataset an uppertorso humanoid robot with 7-DOF arm explored 100 different objects belonging to 20 different categories using 10 behaviors: Look, Crush, Grasp, Hold, Lift, Drop, Poke, Push, Shake and Tap.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We describe the CZ Software Mentions dataset, a new dataset of software mentions in biomedical papers.
1 paper · 0 benchmarks
This dataset comprises video files (converted into tif format) that depict glomerular activation in mice.
1 paper · 0 benchmarks
A collaborative effort between researchers at the Vascular Imaging Lab located at the University of Calgary and the Medical Image Computing Lab located at the University of Campinas (UNICAMP) originated the Calgary Campinas public brain…
1 paper · 0 benchmarks
Calibration data (Observations for intensive stations in the Finnish Archipelago Sea 2006-2015)
Data used for calibrating Finnish Coastal nutrient load model (FICOS).
1 paper · 0 benchmarks
This dataset collects 88,077 numerical samples of call options on Shanghai Stock Exchange from 2015-02 to 2020-07.
1 paper · 0 benchmarks
Calliar is a dataset for Arabic calligraphy.
1 paper · 0 benchmarks
https://zenodo.org/records/15301636
1 paper · 0 benchmarks
Cam-CAN (Cambridge Centre for Ageing and Neuroscience dataset)
The Cambridge Centre for Ageing and Neuroscience (Cam-CAN) is a large-scale collaborative research project at the University of Cambridge, launched in October 2010, with substantial initial funding from the Biotechnology and Biological…
1 paper · 0 benchmarks
57 stock videos from Pexels, predominantly covering road scenes which involve minimal distortion.
1 paper · 0 benchmarks
35 recordings of Candombe music with beat and downbeat annotations.
1 paper · 2 benchmarks
The goal of this project is to present two new datasets that seek to expand the capability of the Learning to See in the Dark Low-light enhancement CNN for the Canon 6D DSLR, and explore how the network performs when modified in various…
1 paper · 2 benchmarks
Consists of eye movements and verbal descriptions recorded synchronously over images.
1 paper · 0 benchmarks
The CapMIT1003 database contains captions and clicks collected for images from the MIT1003 database, for which reference eye scanpath are available.
1 paper · 1 benchmark
CapriDB is a 3D object database for robotics.
1 paper · 0 benchmarks
Capriccio is a sentiment classification dataset on tweets that simulates data drift.
1 paper · 0 benchmarks
This dataset contains Axivity AX3 wrist-worn activity tracker data that were collected from 151 participants in 2014-2016 around the Oxfordshire area.
1 paper · 0 benchmarks
In this dataset we added [Company Name, Car Model, Car Type, Fuel Type, Transmission, Engine (cc), Mileage, Kmsdriven, Buyers, Horsepower (kw), Year Price (Lakhs)]
1 paper · 1 benchmark
A synthetic dataset from an automobile manufacturer datasource.
1 paper · 0 benchmarks
Energy production and carbon intensity datasets for the regions Germany, Great Britain, France (all via the ENTSO-E Transparency Platform) and California (via California ISO) for the entire year 2020 +-10 days.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.