Home › Datasets › language › Chinese

Chinese datasets

archive 2025-07-28

460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 3 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Chinese datasets 97–144 of 460

This dataset contains the traffic data in San Bernardino from July to August in 2016, with 170 detectors on 8 roads with a time interval of 5 minutes.
44 papers · 1 benchmark
The SCUT-CTW1500 dataset contains 1,500 images: 1,000 for training and 500 for testing.
44 papers · 3 benchmarks
AVSpeech is a large-scale audio-visual dataset comprising speech clips with no interfering background signals.
42 papers · 0 benchmarks
BUCC (Building and Using Comparable Corpora)
The BUCC mining task is a shared task on parallel sentence extraction from two monolingual corpora with a subset of them assumed to be parallel, and that has been available since 2016.
42 papers · 4 benchmarks
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems.
38 papers · 0 benchmarks
TextOCR is a dataset to benchmark text recognition on arbitrary shaped scene-text.
37 papers · 0 benchmarks
DialogRE is the first human-annotated dialogue-based relation extraction dataset, containing 1,788 dialogues originating from the complete transcripts of a famous American television situation comedy Friends.
36 papers · 1 benchmark
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
THCHS-30 is a free Chinese speech database THCHS-30 that can be used to build a full-fledged Chinese speech recognition system.
34 papers · 0 benchmarks
WMT 2020 is a collection of datasets used in shared tasks of the Fifth Conference on Machine Translation.
33 papers · 0 benchmarks
MS-CXR (Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing)
The MS-CXR dataset provides 1162 image–sentence pairs of bounding boxes and corresponding phrases, collected across eight different cardiopulmonary radiological findings, with an approximately equal number of pairs for each finding.
32 papers · 0 benchmarks
IGLUE (Image-Grounded Language Understanding Evaluation)
The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded…
31 papers · 0 benchmarks
N-UCLA (Northwestern-UCLA Multiview Action 3D Dataset)
The Multiview 3D event dataset is capture by me and Xiaohan Nie in UCLA.
30 papers · 2 benchmarks
VALUE (Video-And-Language Understanding Evaluation)
VALUE is a Video-And-Language Understanding Evaluation benchmark to test models that are generalizable to diverse tasks, domains, and datasets.
30 papers · 0 benchmarks
MaRVL (Multicultural Reasoning over Vision and Language)
Multicultural Reasoning over Vision and Language (MaRVL) is a dataset based on an ImageNet-style hierarchy representative of many languages and cultures (Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish).
29 papers · 1 benchmark
A human-to-human Chinese dialog dataset (about 10k dialogs, 156k utterances), which contains multiple sequential dialogs for every pair of a recommendation seeker (user) and a recommender (bot).
28 papers · 0 benchmarks
Resume contains eight fine-grained entity categories -score from 74.5% to 86.88%.
28 papers · 1 benchmark
XM 3600 (Crossmodal 3600)
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
27 papers · 1 benchmark
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
XStoryCloze consists of the professionally translated version of the English StoryCloze dataset (Spring 2016 version) to 10 non-English languages.
26 papers · 0 benchmarks
CH-SIMS is a Chinese single- and multimodal sentiment analysis dataset which contains 2,281 refined video segments in the wild with both multimodal and independent unimodal annotations.
25 papers · 1 benchmark
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
CCPD (Chinese City Parking Dataset)
The Chinese City Parking Dataset (CCPD) is a dataset for license plate detection and recognition.
24 papers · 0 benchmarks
BEAT2 (BEAT-SMPLX-FLAME)
We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements.
23 papers · 2 benchmarks
DocUNet (Document Image Unwarping via a Stacked U-Net)
Various documents dataset.
23 papers · 3 benchmarks
MSRA CN NER (MSRA CN NER Dataset)
Simplified Chinese dataset for NER in The Third International Chinese Language Processing Bakeoff (2006), provided by Microsoft Research Asia (MSRA).
23 papers · 3 benchmarks
PointCloud-C is the very first test-suite for point cloud robustness analysis under corruptions.
23 papers · 2 benchmarks
Real 3D-AD is the first point cloud anomaly detection dataset for industrial products.
23 papers · 2 benchmarks
We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
22 papers · 0 benchmarks
KdConv (Knowledge-driven Conversation)
KdConv is a Chinese multi-domain Knowledge-driven Conversation dataset, grounding the topics in multi-turn conversations to knowledge graphs.
22 papers · 0 benchmarks
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
COCO-CN is a bilingual image description dataset enriching MS-COCO with manually written Chinese sentences and tags.
21 papers · 1 benchmark
A team of researchers from Qatar University, Doha, Qatar, and the University of Dhaka, Bangladesh along with their collaborators from Pakistan and Malaysia in collaboration with medical doctors have created a database of chest X-ray images…
21 papers · 0 benchmarks
EPHOIE (phtnsantader@gmail.com)
EPHOIE is a fully-annotated dataset which is the first Chinese benchmark for both text spotting and visual information extraction.
21 papers · 2 benchmarks
MIR-1K (Multimedia Information Retrieval lab, 1000 song clips) is a dataset designed for singing voice separation.
21 papers · 0 benchmarks
MLQE-PE (Multilingual Quality Estimation and Automatic Post-editing Dataset)
The Multilingual Quality Estimation and Automatic Post-editing (MLQE-PE) Dataset is a dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE).
20 papers · 0 benchmarks
PsyQA is a Chinese Dataset for generating long counseling text for mental health support.
20 papers · 0 benchmarks
RCTW-17 (Reading Chinese Text in the Wild)
Features a large-scale dataset with 12,263 annotated images.
20 papers · 0 benchmarks
XFUND (A Multilingual Form Understanding Benchmark)
XFUND is a multilingual form understanding benchmark dataset that includes human-labeled forms with key-value pairs in 7 languages (Chinese, Japanese, Spanish, French, Italian, German, Portuguese).
20 papers · 0 benchmarks
The iKala dataset is a singing voice separation dataset that comprises of 252 30-second excerpts sampled from 206 iKala songs (plus 100 hidden excerpts reserved for MIREX data mining contest).
20 papers · 1 benchmark
Ten years (2008-2018) ChFinAnn documents and human-summarized event knowledge bases to conduct the DS-based event labeling.
19 papers · 1 benchmark
Chaoyang dataset contains 1111 normal, 842 serrated, 1404 adenocarcinoma, 664 adenoma, and 705 normal, 321 serrated, 840 adenocarcinoma, 273 adenoma samples for training and testing, respectively.
19 papers · 2 benchmarks
Contains data from three platforms, i.e., synthetic drones, satellites and ground cameras of 1,652 university buildings around the world.
19 papers · 2 benchmarks
The first parallel corpus composed from United Nations documents published by the original data creator.
18 papers · 0 benchmarks
Weibo21 is a benchmark of fake news dataset for multi-domain fake news detection (MFND) with domain label annotated, which consists of 4,488 fake news and 4,640 real news from 9 different domains.
18 papers · 0 benchmarks
xSID (Cross-lingual Slot and Intent Detection)
xSID, a new evaluation benchmark for cross-lingual (X) Slot and Intent Detection in 13 languages from 6 language families, including a very low-resource dialect, covering Arabic (ar), Chinese (zh), Danish (da), Dutch (nl), English (en),…
18 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.