Home › Datasets › language › Chinese
Chinese datasets
archive 2025-07-28
460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 4 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Chinese datasets 145–192 of 460
CASIA-HWDB is a dataset for handwritten Chinese character recognition.
17 papers · 0 benchmarks
LCCC (Large-scale Cleaned Chinese Conversation corpus)
Contains a base version (6.8million dialogues) and a large version (12.0 million dialogues).
17 papers · 0 benchmarks
MathBench (MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark)
MathBench is an All in One math dataset for language model evaluation, with: A Sophisticated Five-Stage Difficulty Mechanism: Unlike the usual mathematical datasets that can only evaluate a single difficulty level or have a mix of unclear…
16 papers · 0 benchmarks
OntoNotes Release 4.0 contains the content of earlier releases -- OntoNotes Release 1.0 LDC2007T21, OntoNotes Release 2.0 LDC2008T04 and OntoNotes Release 3.0 LDC2009T24 -- and adds newswire, broadcast news, broadcast conversation and web…
16 papers · 1 benchmark
Traffic (Traffic Flow Forecasting Data Set)
Abstract: The task for this dataset is to forecast the spatio-temporal traffic volume based on the historical traffic volume and other features in neighboring locations.
16 papers · 2 benchmarks
Wukong is a large-scale Chinese cross-modal dataset for benchmarking different multi-modal pre-training methods to facilitate the Vision-Language Pre-training (VLP).
16 papers · 0 benchmarks
CPED (Chinese Personalized and Emotional Dialogue)
We construct a dataset named CPED from 40 Chinese TV shows.
15 papers · 3 benchmarks
Data was collected for normal bearings, single-point drive end and fan end defects.
15 papers · 1 benchmark
MuCGEC (Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction)
MuCGEC is a multi-reference multi-source evaluation dataset for Chinese Grammatical Error Correction (CGEC), consisting of 7,063 sentences collected from three different Chinese-as-a-Second-Language (CSL) learner sources.
15 papers · 1 benchmark
DuLeMon (Baidu Long-term Memory Conversation)
DuLeMon is a large-scale Chinese Long-term Memory Conversation dataset, which simulates long-term memory conversations and focuses on the ability to actively construct and utilize the user's and the bot's persona in a long-term interaction.
14 papers · 0 benchmarks
KUAKE-QIC (Query Intent Classification Dataset)
KUAKE Query Intent Classification, a dataset for intent classification, is used for the KUAKE-QIC task.
14 papers · 1 benchmark
SIBR (SIBR Dataset for VIE in the Wild)
SIBR是面向自然场景视觉信息抽取的数据集。 1)SIBR总的有1000张图片,400张测试,600张训练,包括中文、英文两种语言。 2)包含images.zip、label.zip、train.txt、test.txt四个文件,images.zip、label.zip中包含所有图片和标签,通过train.txt和test.txt区分训练和测试。…
14 papers · 1 benchmark
CirCor DigiScope is currently the largest pediatric heart sound dataset.
13 papers · 2 benchmarks
JEC-QA is a LQA (Legal Question Answering) dataset collected from the National Judicial Examination of China.
13 papers · 0 benchmarks
PersonalDialog is a large-scale multi-turn dialogue dataset containing various traits from a large number of speakers.
13 papers · 0 benchmarks
The SARDet-100K dataset encompasses a total of 116,598 images, and 245,653 instances distributed across six categories: Aircraft, Ship, Car, Bridge, Tank, and Harbor.
13 papers · 1 benchmark
Chinese Few-shot Learning Evaluation Benchmark (FewCLUE) is a comprehensive small sample evaluation benchmark in Chinese.
12 papers · 5 benchmarks
HRSC2016 (High resolution ship collections 2016)
High-resolution ship collections 2016 (HRSC2016) is a data set used for scientific research.
12 papers · 1 benchmark
The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.
12 papers · 0 benchmarks
(1) provide financial news for each specific stock.
11 papers · 2 benchmarks
BSTC (Baidu Speech Translation Corpus)
BSTC (Baidu Speech Translation Corpus) is a large-scale dataset for automatic simultaneous interpretation.
11 papers · 0 benchmarks
CMRC 2017 (Chinese Machine Reading Comprehension 2017)
Contains two different types: cloze-style reading comprehension and user query reading comprehension, associated with large-scale training data as well as human-annotated validation and hidden test set.
11 papers · 0 benchmarks
ICVL is a hyperspectral image dataset, collected by "Sparse Recovery of Hyperspectral Signal from Natural RGB Images" The database images were acquired using a Specim PS Kappa DX4 hyperspectral camera and a rotary stage for spatial…
11 papers · 0 benchmarks
Kaggle EyePACS (Kaggle EyePACS. Diabetic Retinopathy Detection Identify signs of diabetic retinopathy in eye images)
Diabetic retinopathy is the leading cause of blindness in the working-age population of the developed world.
11 papers · 1 benchmark
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
11 papers · 2 benchmarks
A large-scale non-homogeneous remote sensing image dehazing dataset
11 papers · 1 benchmark
Synbols is a dataset generator designed for probing the behavior of learning algorithms.
11 papers · 0 benchmarks
CJRC (Chinese judicial reading comprehension)
The Chinese judicial reading comprehension (CJRC) dataset contains approximately 10K documents and almost 50K questions with answers.
10 papers · 0 benchmarks
CMeEE (Chinese Medical Named Entity Recognition Dataset)
Chinese Medical Named Entity Recognition, a dataset first released in CHIP20204, is used for CMeEE task.
10 papers · 1 benchmark
ChatHaruhi (ChatHaruhi: Reviving Anime Character in Reality via Large Language Model)
ChatHaruhi is a dataset covering 32 Chinese / English TV / anime characters with over 54k simulated dialogues.
10 papers · 0 benchmarks
FM-IQA (Freestyle Multilingual Image Question Answering)
FM-IQA is a question-answering dataset containing over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations.
10 papers · 0 benchmarks
Dataset Introduction In this work, we introduce the In-Diagram Logic (InDL) dataset, an innovative resource crafted to rigorously evaluate the logic interpretation abilities of deep learning models.
10 papers · 1 benchmark
- NeoRL is a collection of environments and datasets for offline reinforcement learning with a special focus on real-world applications.
10 papers · 0 benchmarks
BiToD is a bilingual multi-domain dataset for end-to-end task-oriented dialogue modeling.
9 papers · 0 benchmarks
CODEBRIM (COncrete DEfect BRidge IMage Dataset)
Dataset for multi-target classification of five commonly appearing concrete defects.
9 papers · 0 benchmarks
E-KAR (Benchmark for Explainable Knowledge-intensive Analogical Reasoning)
The ability to recognize analogies is fundamental to human cognition.
9 papers · 0 benchmarks
KaMed is a knowledge-aware medical dialogue dataset, which contains over 60,000 medical dialogue sessions with 5,682 entities (such as Asthma and Atropine).
9 papers · 0 benchmarks
MMCU (Measuring Massive Multitask Chinese Understanding)
We propose a test to measure the multitask accuracy of large Chinese language models.
9 papers · 0 benchmarks
WebCPM is a Chinese LFQA dataset.
9 papers · 0 benchmarks
Abstract Objective This article summarizes the preparation, organization, evaluation, and results of Track 2 of the 2018 National NLP Clinical Challenges shared task.
8 papers · 0 benchmarks
CHIP-CDN (Clinical Diagnosis Normalization Dataset)
CHIP Clinical Diagnosis Normalization, a dataset that aims to standardize the terms from the final diagnoses of Chinese electronic medical records, is used for the CHIP-CDN task.
8 papers · 0 benchmarks
CHIP-STS (Semantic Textual Similarity Dataset)
CHIP Semantic Textual Similarity, a dataset for sentence similarity in the non-i.i.d.
8 papers · 1 benchmark
CMRC 2019 (Chinese Machine Reading Comprehension 2019)
CMRC 2019 is a Chinese Machine Reading Comprehension dataset that was used in The Third Evaluation Workshop on Chinese Machine Reading Comprehension.
8 papers · 0 benchmarks
Emilia Dataset (An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation)
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data.
8 papers · 0 benchmarks
RFUND (Revised FUNSD and XFUND)
RFUND is a relabeled version of FUNSD and XFUND datasets, tackling the following issues in their original annotations: 1.
8 papers · 0 benchmarks
RRS (Restoration-200k for Response Selection)
| | Train | Validation | Test | Ranking Test | | --------- | ----- | ---------- | ------- | ------------ | | size | 0.4M | 50K | 5K | 800 | | pos:neg | 1:1 | 1:9 | 1.2:8.8 | - | | avg turns | 5.0 | 5.0 | 5.0 | 5.0 | Ranking test set…
8 papers · 1 benchmark
News translation is a recurring WMT task.
8 papers · 0 benchmarks
ZEB (Zero-shot Evaluation Benchmark)
A evaluation benchmark ZEB for image matching by merging 8 real-world datasets and 4 simulated datasets with diverse image resolutions, scene conditions and view points.
8 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.