Home › Datasets › language › Chinese
Chinese datasets
archive 2025-07-28
460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 5 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Chinese datasets 193–240 of 460
CASME II (Chinese Academy of Sciences Micro-Expression II)
The Chinese Academy of Sciences Micro-Expression dataset (CASME II) consists of 255 videos, elicited from 26 participants.
7 papers · 1 benchmark
Demetr is a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological…
7 papers · 0 benchmarks
KUAKE Query-Query Relevance, a dataset used to evaluate the relevance of the content expressed in two queries, is used for the KUAKE-QQR task.
7 papers · 1 benchmark
KUAKE Query Title Relevance, a dataset used to estimate the relevance of the title of a query document, is used for the KUAKE-QTR task.
7 papers · 1 benchmark
LEVEN (Legal Event Detection Dataset)
Overview LEVEN is the largest Legal Event Detection dataset as well as the largest Chinese Event Detection dataset.
7 papers · 0 benchmarks
The MedDialog dataset (Chinese) contains conversations (in Chinese) between doctors and patients.
7 papers · 0 benchmarks
A multivariate spatio-temporal benchmark dataset for meteorological forecasting based on real-time observation data from ground weather stations.
7 papers · 16 benchmarks
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese code-switching corpus collected in Hong Kong.
6 papers · 0 benchmarks
BMELD is a bilingual (English-Chinese) dialogue corpus for Neural chat translation.
6 papers · 0 benchmarks
BiPaR is a manually annotated bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support monolingual, multilingual and cross-lingual reading comprehension on novels.
6 papers · 0 benchmarks
CAIS (Chinese Artificial Intelligence Speakers)
We collect utterances from the Chinese Artificial Intelligence Speakers (CAIS), and annotate them with slot tags and intent labels.
6 papers · 2 benchmarks
CLUECorpus2020 is a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation.
6 papers · 0 benchmarks
ChineseFoodNet aims to automatically recognizing pictured Chinese dishes.
6 papers · 0 benchmarks
Concepticon (Concepticon. A Resource for the Linking of Concept Lists)
This resource, our Concepticon, links concept labels from different conceptlists to concept sets.
6 papers · 0 benchmarks
This is a synthetic dataset for defect detection on textured surfaces.
6 papers · 1 benchmark
Cant (also known as doublespeak, cryptolect, argot, anti-language or secret language) is important for understanding advertising, comedies and dog-whistle politics.
6 papers · 0 benchmarks
This dataset is a new benchmark, grounded in real-world usages is developed to support more authentic and comprehensive evaluation of image editing models.
6 papers · 2 benchmarks
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
We introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency.
6 papers · 0 benchmarks
LCQMC (Large-scale Chinese Question Matching Corpus)
LCQMC is a large-scale Chinese question matching corpus.
6 papers · 0 benchmarks
LIS (low-light instance segmentation)
To reveal and systematically investigate the effectiveness of the proposed method in the real world, a real low-light image dataset for instance segmentation is necessary and urgently needed.
6 papers · 0 benchmarks
Lyra is a dataset for code generation that consists on Python code with embedded SQL.
6 papers · 0 benchmarks
SWSR (Sina Weibo Sexism Review)
The Sina Weibo Sexism Review (SWSR) dataset is a dataset to research online sexism in Chinese.
6 papers · 0 benchmarks
We present a further analysis of visual modality incompleteness, benchmarking latest MMEA models on our proposed dataset MMEA-UMVM.
6 papers · 3 benchmarks
Ultra-high definition benchmark (UHDBench) includes 2293 images at 2k resolution sourced from the ground-truth test sets of HRSOD, LIU4k, UAVid, UHDM, and UHRSD.
6 papers · 1 benchmark
XImageNet-12 (XIMAGENET-12: An Explainable AI Benchmark Dataset for Model Robustness Evaluation)
Enlarge the dataset to understand how image background effect the Computer Vision ML model.
6 papers · 1 benchmark
XL-BEL is a benchmark for cross-lingual biomedical entity linking (XL-BEL).
6 papers · 0 benchmarks
XQA is a data which consists of a total amount of 90k question-answer pairs in nine languages for cross-lingual open-domain question answering.
6 papers · 0 benchmarks
BIWI 3D corpus comprises a total of 1109 sentences uttered by 14 native English speakers (6 males and 8 females).
5 papers · 1 benchmark
CCPM (Chinese Classical Poetry Matching)
Introduction CCPM is a large Chinese classical poetry matching dataset that can be used for poetry matching, understanding and translation.
5 papers · 0 benchmarks
Chinese dataset on COVID-19 misinformation.
5 papers · 0 benchmarks
Chinese Text in the Wild is a dataset of Chinese text with about 1 million Chinese characters from 3850 unique ones annotated by experts in over 30000 street view images.
5 papers · 0 benchmarks
DurLAR (A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery)
DurLAR is a high-fidelity 128-channel 3D LiDAR dataset with panoramic ambient (near infrared) and reflectivity imagery for multi-modal autonomous driving applications.
5 papers · 0 benchmarks
We collect, organize and open-source the large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples.
5 papers · 0 benchmarks
We propose a novel long-context benchmark, 🐉 Loong, aligning with realistic scenarios through extended multi-document question answering (QA).
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
MATINF (Maternal and Infant Dataset)
Maternal and Infant (MATINF) Dataset is a large-scale dataset jointly labeled for classification, question answering and summarization in the domain of maternity and baby caring in Chinese.
5 papers · 0 benchmarks
The Messidor database has been established to facilitate studies on computer-assisted diagnoses of diabetic retinopathy.
5 papers · 0 benchmarks
- A large scale Chinese multi-modal dialogue corpus (120.84K dialogues and 198.82 K images).
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
RRS Ranking Test (Restoration-200k for Response Selection with Ranking Test Set)
| | Train | Validation | Test | Ranking Test | | --------- | ----- | ---------- | ------- | ------------ | | size | 0.4M | 50K | 5K | 800 | | pos:neg | 1:1 | 1:9 | 1.2:8.8 | - | | avg turns | 5.0 | 5.0 | 5.0 | 5.0 | Ranking test set…
5 papers · 1 benchmark
RealMAN (A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization)
The Audio Signal and Information Processing Lab at Westlake University, in collaboration with AISHELL, has released the Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset, which provides annotated multi-channel…
5 papers · 2 benchmarks
RiSAWOZ is a large-scale multi-domain Chinese Wizard-of-Oz dataset with Rich Semantic Annotations.
5 papers · 0 benchmarks
S.MID (SeMantic InDustry)
SeMantic InDustry (S.MID) is a dataset designed to advance the field of LiDAR semantic segmentation, specifically for robotic applications and large-scale industrial scene.
5 papers · 1 benchmark
YACLC (Yet Another Chinese Learner Corpus)
YACLC is a large scale, multidimensional annotated Chinese learner corpus.
5 papers · 0 benchmarks
mTVR is a large-scale multilingual video moment retrieval dataset, containing 218K English and Chinese queries from 21.8K TV show video clips.
5 papers · 0 benchmarks
CUGE is a Chinese Language Understanding and Generation Evaluation benchmark with the following features: (1) Hierarchical benchmark framework, where datasets are principally selected and organized with a language capability-task-dataset…
4 papers · 0 benchmarks
The ChineseLP dataset contains 411 vehicle images (mostly of passenger cars) with Chinese license plates (LPs).
4 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.