Home › Datasets › language › Chinese

Chinese datasets

archive 2025-07-28

460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 4 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Chinese datasets 145–192 of 460

CASIA-HWDB is a dataset for handwritten Chinese character recognition.
17 papers · 0 benchmarks
LCCC (Large-scale Cleaned Chinese Conversation corpus)
Contains a base version (6.8million dialogues) and a large version (12.0 million dialogues).
17 papers · 0 benchmarks
MathBench (MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark)
MathBench is an All in One math dataset for language model evaluation, with: A Sophisticated Five-Stage Difficulty Mechanism: Unlike the usual mathematical datasets that can only evaluate a single difficulty level or have a mix of unclear…
16 papers · 0 benchmarks
OntoNotes 4.0 (OntoNotes Release 4.0)
OntoNotes Release 4.0 contains the content of earlier releases -- OntoNotes Release 1.0 LDC2007T21, OntoNotes Release 2.0 LDC2008T04 and OntoNotes Release 3.0 LDC2009T24 -- and adds newswire, broadcast news, broadcast conversation and web…
16 papers · 1 benchmark
Traffic (Traffic Flow Forecasting Data Set)
Abstract: The task for this dataset is to forecast the spatio-temporal traffic volume based on the historical traffic volume and other features in neighboring locations.
16 papers · 2 benchmarks
Wukong is a large-scale Chinese cross-modal dataset for benchmarking different multi-modal pre-training methods to facilitate the Vision-Language Pre-training (VLP).
16 papers · 0 benchmarks
CPED (Chinese Personalized and Emotional Dialogue)
We construct a dataset named CPED from 40 Chinese TV shows.
15 papers · 3 benchmarks
Data was collected for normal bearings, single-point drive end and fan end defects.
15 papers · 1 benchmark
MuCGEC (Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction)
MuCGEC is a multi-reference multi-source evaluation dataset for Chinese Grammatical Error Correction (CGEC), consisting of 7,063 sentences collected from three different Chinese-as-a-Second-Language (CSL) learner sources.
15 papers · 1 benchmark
DuLeMon (Baidu Long-term Memory Conversation)
DuLeMon is a large-scale Chinese Long-term Memory Conversation dataset, which simulates long-term memory conversations and focuses on the ability to actively construct and utilize the user's and the bot's persona in a long-term interaction.
14 papers · 0 benchmarks
KUAKE-QIC (Query Intent Classification Dataset)
KUAKE Query Intent Classification, a dataset for intent classification, is used for the KUAKE-QIC task.
14 papers · 1 benchmark
SIBR (SIBR Dataset for VIE in the Wild)
SIBR是面向自然场景视觉信息抽取的数据集。 1)SIBR总的有1000张图片,400张测试,600张训练,包括中文、英文两种语言。 2)包含images.zip、label.zip、train.txt、test.txt四个文件,images.zip、label.zip中包含所有图片和标签,通过train.txt和test.txt区分训练和测试。…
14 papers · 1 benchmark
CirCor DigiScope is currently the largest pediatric heart sound dataset.
13 papers · 2 benchmarks
JEC-QA is a LQA (Legal Question Answering) dataset collected from the National Judicial Examination of China.
13 papers · 0 benchmarks
PersonalDialog is a large-scale multi-turn dialogue dataset containing various traits from a large number of speakers.
13 papers · 0 benchmarks
The SARDet-100K dataset encompasses a total of 116,598 images, and 245,653 instances distributed across six categories: Aircraft, Ship, Car, Bridge, Tank, and Harbor.
13 papers · 1 benchmark
Chinese Few-shot Learning Evaluation Benchmark (FewCLUE) is a comprehensive small sample evaluation benchmark in Chinese.
12 papers · 5 benchmarks
HRSC2016 (High resolution ship collections 2016)
High-resolution ship collections 2016 (HRSC2016) is a data set used for scientific research.
12 papers · 1 benchmark
The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.
12 papers · 0 benchmarks
(1) provide financial news for each specific stock.
11 papers · 2 benchmarks
BSTC (Baidu Speech Translation Corpus)
BSTC (Baidu Speech Translation Corpus) is a large-scale dataset for automatic simultaneous interpretation.
11 papers · 0 benchmarks
CMRC 2017 (Chinese Machine Reading Comprehension 2017)
Contains two different types: cloze-style reading comprehension and user query reading comprehension, associated with large-scale training data as well as human-annotated validation and hidden test set.
11 papers · 0 benchmarks
ICVL is a hyperspectral image dataset, collected by "Sparse Recovery of Hyperspectral Signal from Natural RGB Images" The database images were acquired using a Specim PS Kappa DX4 hyperspectral camera and a rotary stage for spatial…
11 papers · 0 benchmarks
Kaggle EyePACS (Kaggle EyePACS. Diabetic Retinopathy Detection Identify signs of diabetic retinopathy in eye images)
Diabetic retinopathy is the leading cause of blindness in the working-age population of the developed world.
11 papers · 1 benchmark
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
11 papers · 2 benchmarks
A large-scale non-homogeneous remote sensing image dehazing dataset
11 papers · 1 benchmark
Synbols is a dataset generator designed for probing the behavior of learning algorithms.
11 papers · 0 benchmarks
CJRC (Chinese judicial reading comprehension)
The Chinese judicial reading comprehension (CJRC) dataset contains approximately 10K documents and almost 50K questions with answers.
10 papers · 0 benchmarks
CMeEE (Chinese Medical Named Entity Recognition Dataset)
Chinese Medical Named Entity Recognition, a dataset first released in CHIP20204, is used for CMeEE task.
10 papers · 1 benchmark
ChatHaruhi (ChatHaruhi: Reviving Anime Character in Reality via Large Language Model)
ChatHaruhi is a dataset covering 32 Chinese / English TV / anime characters with over 54k simulated dialogues.
10 papers · 0 benchmarks
FM-IQA (Freestyle Multilingual Image Question Answering)
FM-IQA is a question-answering dataset containing over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations.
10 papers · 0 benchmarks
InDL (In-Diagram Logic)
Dataset Introduction In this work, we introduce the In-Diagram Logic (InDL) dataset, an innovative resource crafted to rigorously evaluate the logic interpretation abilities of deep learning models.
10 papers · 1 benchmark
- NeoRL is a collection of environments and datasets for offline reinforcement learning with a special focus on real-world applications.
10 papers · 0 benchmarks
BiToD is a bilingual multi-domain dataset for end-to-end task-oriented dialogue modeling.
9 papers · 0 benchmarks
CODEBRIM (COncrete DEfect BRidge IMage Dataset)
Dataset for multi-target classification of five commonly appearing concrete defects.
9 papers · 0 benchmarks
E-KAR (Benchmark for Explainable Knowledge-intensive Analogical Reasoning)
The ability to recognize analogies is fundamental to human cognition.
9 papers · 0 benchmarks
KaMed is a knowledge-aware medical dialogue dataset, which contains over 60,000 medical dialogue sessions with 5,682 entities (such as Asthma and Atropine).
9 papers · 0 benchmarks
MMCU (Measuring Massive Multitask Chinese Understanding)
We propose a test to measure the multitask accuracy of large Chinese language models.
9 papers · 0 benchmarks
WebCPM is a Chinese LFQA dataset.
9 papers · 0 benchmarks
2018 n2c2 (Track 2) - Adverse Drug Events and Medication Extraction (2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records)
Abstract Objective This article summarizes the preparation, organization, evaluation, and results of Track 2 of the 2018 National NLP Clinical Challenges shared task.
8 papers · 0 benchmarks
CHIP-CDN (Clinical Diagnosis Normalization Dataset)
CHIP Clinical Diagnosis Normalization, a dataset that aims to standardize the terms from the final diagnoses of Chinese electronic medical records, is used for the CHIP-CDN task.
8 papers · 0 benchmarks
CHIP-STS (Semantic Textual Similarity Dataset)
CHIP Semantic Textual Similarity, a dataset for sentence similarity in the non-i.i.d.
8 papers · 1 benchmark
CMRC 2019 (Chinese Machine Reading Comprehension 2019)
CMRC 2019 is a Chinese Machine Reading Comprehension dataset that was used in The Third Evaluation Workshop on Chinese Machine Reading Comprehension.
8 papers · 0 benchmarks
Emilia Dataset (An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation)
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data.
8 papers · 0 benchmarks
RFUND (Revised FUNSD and XFUND)
RFUND is a relabeled version of FUNSD and XFUND datasets, tackling the following issues in their original annotations: 1.
8 papers · 0 benchmarks
RRS (Restoration-200k for Response Selection)
| | Train | Validation | Test | Ranking Test | | --------- | ----- | ---------- | ------- | ------------ | | size | 0.4M | 50K | 5K | 800 | | pos:neg | 1:1 | 1:9 | 1.2:8.8 | - | | avg turns | 5.0 | 5.0 | 5.0 | 5.0 | Ranking test set…
8 papers · 1 benchmark
WMT 2018 News (WMT 2018 News Translation Task)
News translation is a recurring WMT task.
8 papers · 0 benchmarks
ZEB (Zero-shot Evaluation Benchmark)
A evaluation benchmark ZEB for image matching by merging 8 real-world datasets and 4 simulated datasets with diverse image resolutions, scene conditions and view points.
8 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.