Home › Datasets › language › Chinese
Chinese datasets
archive 2025-07-28
460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 2 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Chinese datasets 49–96 of 460
A repository that contains political events with a specific timestamp.
113 papers · 2 benchmarks
GlaS (Gland Segmentation in Colon Histology Images Challenge)
The dataset used in this challenge consists of 165 images derived from 16 H&E stained histological sections of stage T3 or T42 colorectal adenocarcinoma.
112 papers · 1 benchmark
MGSM (Multilingual Grade School Math)
Multilingual Grade School Math Benchmark (MGSM) is a benchmark of grade-school math problems.
107 papers · 1 benchmark
Comprises 11 hand gesture categories from 29 subjects under 3 illumination conditions.
103 papers · 6 benchmarks
The Synthetic Rain Datasets consists of 13,712 clean-rain image pairs gathered from multiple datasets (Rain14000, Rain1800, Rain800, Rain12).
102 papers · 6 benchmarks
The CrowdPose dataset contains about 20,000 images and a total of 80,000 human poses with 14 labeled keypoints.
99 papers · 2 benchmarks
The LUNA16 (LUng Nodule Analysis) dataset is a dataset for lung segmentation.
99 papers · 0 benchmarks
Math23K (Math23K for Math Word Problem Solving)
Math23K is a dataset created for math word problem solving, contains 23, 162 Chinese problems crawled from the Internet.
95 papers · 1 benchmark
ASPEC (Asian Scientific Paper Excerpt Corpus)
ASPEC, Asian Scientific Paper Excerpt Corpus, is constructed by the Japan Science and Technology Agency (JST) in collaboration with the National Institute of Information and Communications Technology (NICT).
87 papers · 0 benchmarks
PPMI (Parkinson’s Progression Markers Initiative)
The Parkinson’s Progression Markers Initiative (PPMI) dataset originates from an observational clinical and longitudinal study comprising evaluations of people with Parkinson’s disease (PD), those people with high risk, and those who are…
87 papers · 3 benchmarks
Douban (Douban Conversation Corpus)
We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
81 papers · 4 benchmarks
MPIIGaze is a dataset for appearance-based gaze estimation in the wild.
78 papers · 2 benchmarks
CoNaLa (CMU CoNaLa, the Code/Natural Language Challenge)
The CMU CoNaLa, the Code/Natural Language Challenge dataset is a joint project from the Carnegie Mellon University NeuLab and Strudel labs.
77 papers · 1 benchmark
QM9 provides quantum chemical properties (at DFT level) for a relevant, consistent, and comprehensive chemical space of small organic molecules.
76 papers · 9 benchmarks
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
The shared task of CoNLL-2002 concerns language-independent named entity recognition.
70 papers · 3 benchmarks
MuPoTS-3D (Multiperson Pose Test Set in 3DMulti-person Pose estimation Test Set in 3D)
MuPoTs-3D (Multi-person Pose estimation Test Set in 3D) is a dataset for pose estimation composed of more than 8,000 frames from 20 real-world scenes with up to three subjects.
70 papers · 3 benchmarks
CMRC (Chinese Machine Reading Comprehension)
CMRC is a dataset is annotated by human experts with near 20,000 questions as well as a challenging set which is composed of the questions that need reasoning over multiple clues.
69 papers · 0 benchmarks
CN-Celeb is a large-scale speaker recognition dataset collected in the wild'.
68 papers · 1 benchmark
CoVoST2 (Common Voice Speech-To-Text 2)
End-to-end speech-to-text translation (ST) has recently witnessed an increased interest given its system simplicity, lower inference latency and less compounding errors compared to cascaded ST (i.e.
66 papers · 0 benchmarks
DBP15k contains four language-specific KGs that are respectively extracted from English (En), Chinese (Zh), French (Fr) and Japanese (Ja) DBpedia, each of which contains around 65k-106k entities.
65 papers · 3 benchmarks
DuReader is a large-scale open-domain Chinese machine reading comprehension dataset.
65 papers · 1 benchmark
FakeAVCeleb is a novel Audio-Video Deepfake dataset that not only contains deepfake videos but respective synthesized cloned audios as well.
65 papers · 2 benchmarks
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
CSL-Daily (Chinese Sign Language Corpus) is a large-scale continuous SLT dataset.
63 papers · 2 benchmarks
ESD (Emotional Speech Database)
ESD is an Emotional Speech Database for voice conversion research.
63 papers · 0 benchmarks
SLAKE is an English-Chinese bilingual dataset consisting of 642 images and 14,028 question-answer pairs for training and testing Med-VQA systems.
63 papers · 0 benchmarks
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 dialogues with corresponding manually labeled summaries and topics.
62 papers · 2 benchmarks
The LIP (Look into Person) dataset is a large-scale dataset focusing on semantic understanding of a person.
61 papers · 1 benchmark
LCSTS is a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public.
58 papers · 2 benchmarks
BEAT (Body-Expression-Audio-Text)
BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations.
56 papers · 1 benchmark
PA-100K is a recent-proposed large pedestrian attribute dataset, with 100,000 images in total collected from outdoor surveillance cameras.
56 papers · 1 benchmark
Emotion-cause pair extraction (ECPE) aims to extract the potential pairs of emotions and corresponding causes in a document.
55 papers · 0 benchmarks
C3 is a free-form multiple-Choice Chinese machine reading Comprehension dataset.
54 papers · 0 benchmarks
CVC-ClinicDB is an open-access dataset of 612 images with a resolution of 384×288 from 31 colonoscopy sequences.It is used for medical image segmentation, in particular polyp detection in colonoscopy videos.
54 papers · 1 benchmark
The Epinions dataset is built form a who-trust-whom online social network of a general consumer review site Epinions.com.
54 papers · 2 benchmarks
DRCD (Delta Reading Comprehension Dataset)
Delta Reading Comprehension Dataset (DRCD) is an open domain traditional Chinese machine reading comprehension (MRC) dataset.
53 papers · 0 benchmarks
Large language models (LLMs), after being aligned with vision models and integrated into vision-language models (VLMs), can bring impressive improvement in image reasoning tasks.
53 papers · 1 benchmark
MLDoc (Multilingual Document Classification Corpus)
Multilingual Document Classification Corpus (MLDoc) is a cross-lingual document classification dataset covering English, German, French, Spanish, Italian, Russian, Japanese and Chinese.
53 papers · 8 benchmarks
The Weibo NER dataset is a Chinese Named Entity Recognition dataset drawn from the social media website Sina Weibo.
52 papers · 2 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
MKQA (Multilingual Knowledge Questions and Answers)
Multilingual Knowledge Questions and Answers (MKQA) is an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically diverse languages (260k question-answer pairs in total).
49 papers · 0 benchmarks
CityFlow is a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 10 intersections, with the longest distance between two simultaneous cameras being 2.5 km.
47 papers · 1 benchmark
AliMeeting (Multi-Channel Multi-Party Meeting Transcription Challenge)
AliMeeting corpus consists of 120 hours of recorded Mandarin meeting data, including far-field data collected by 8-channel microphone array as well as near-field data collected by headset microphone.
46 papers · 1 benchmark
LiTS17 (Liver Tumor Segmentation Challenge 2017)
LiTS17 is a liver tumor segmentation benchmark.
45 papers · 3 benchmarks
OCNLI (Original Chinese Natural Language Inference)
OCNLI stands for Original Chinese Natural Language Inference.
44 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.