Home › Datasets › language › Chinese

Chinese datasets

archive 2025-07-28

460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 8 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Chinese datasets 337–384 of 460

CHORD (CHOrus Recognition Dataset)
CHORD is the first chorus recognition dataset containing 627 songs for public use.
1 paper · 0 benchmarks
CLPD (China License Plate Dataset)
The CLPD dataset comprises 1200 images that encompass various regions within mainland China.
1 paper · 0 benchmarks
CMACD (Chinese Multi-label Affective Computing Dataset)
This study collected data from the major social media platform Weibo, screening 11,338 valid users from over 50,000 individuals with diverse MBTI personality labels and acquiring 566,900 posts along with the user MBTI personality tags.
1 paper · 0 benchmarks
CMeIE (Chinese Medical Information Extraction Dataset)
Chinese Medical Information Extraction, a dataset that is also released in CHIP2020, is used for CMeIE task.
1 paper · 1 benchmark
Contains a dataset of 241 Chinese dishes with 191,811 images.
1 paper · 0 benchmarks
CNFOOD-241 Contains a dataset of 241 Chinese dishes with 191,811 images.
1 paper · 1 benchmark
We present CSL, a large-scale Chinese Scientific Literature dataset, which contains the titles, abstracts, keywords and academic fields of 396,209 papers.
1 paper · 0 benchmarks
A large-scale gloss-free sign language translation dataset with 1,985 hours of videos, approximately 86 times larger than the previous CSL-Daily dataset.
1 paper · 0 benchmarks
CSPRD (Chinese Stock Policy Retrieval Dataset)
The Chinese Stock Policy Retrieval Dataset (CSPRD) contains a Chinese policy corpus of 10,002 articles and 709 prospectus examples from 545 companies listed on China’s Science and Technology Innovation Board (STAR Market).
1 paper · 0 benchmarks
A synthetic dataset from an automobile manufacturer datasource.
1 paper · 0 benchmarks
ChCatExt (Chinese Catalog Extraction Dataset)
ChCatExt is composed of BidAnn (bid announcement), FinAnn (financial announcement) and CreRat (credit rating report).
1 paper · 1 benchmark
ChinaOpen is a new video dataset targeted at open-world multimodal learning, with raw data gathered from Bilibili, a popular Chinese video-sharing website.
1 paper · 1 benchmark
Large-scale Chinese legal dataset for judgment prediction.
1 paper · 0 benchmarks
Chinese Literature NER RE is a Discourse-Level Named Entity Recognition and Relation Extraction Dataset for Chinese Literature Text.
1 paper · 0 benchmarks
The Chinese Traditional Painting dataset for style transfer contains 1000 content images and 100 style images.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Conic10K is an open-ended math problem dataset on conic sections in Chinese senior high school education.
1 paper · 0 benchmarks
ConvSumX is a cross-lingual conversation summarization benchmark, through a new annotation schema that explicitly considers source input context.
1 paper · 0 benchmarks
DIGITal (Digitally Generated Numerals)
Digitally Generated Numerals (DIGITal) Description The Digitally Generated Numerals (DIGITal) dataset consists of 100,000 image pairs representing digits from 0 to 9.
1 paper · 0 benchmarks
dataset in WWW 2019 "DPLink: User Identity Linkage via Deep Neural Network From Heterogeneous Mobility Data".
1 paper · 0 benchmarks
This data is for the Mis2-KDD 2021 under review paper: Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People’s Republic of China We present our dataset that focuses on propaganda techniques in Mandarin…
1 paper · 1 benchmark
DiaKG is a high-quality Chinese dataset for Diabetes knowledge graph.
1 paper · 0 benchmarks
EUCA dataset description Associated Paper: EUCA: the End-User-Centered Explainable AI Framework Authors: Weina Jin, Jianyu Fan, Diane Gromala, Philippe Pasquier, Ghassan Hamarneh Introduction: EUCA dataset is for modelling personalized or…
1 paper · 0 benchmarks
FCGEC (FCGEC: Fine-Grained Corpus for Chinese Grammatical Error Correction)
a fine-grained corpus to detect, identify and correct the chinese grammatical errors.
1 paper · 1 benchmark
FGraDA (Fine-Grained Domain Adaptation Dataset)
Previous research for adapting a general neural machine translation (NMT) model into a specific domain usually neglects the diversity in translation within the same domain, which is a core problem for domain adaptation in real- world…
1 paper · 0 benchmarks
ForPKG (https://github.com/luozhongze/ForPKG)
A policy knowledge graph can provide decision support for tasks such as project compliance, policy analysis, and intelligent question answering, and can also serve as an external knowledge base to assist the reasoning process of related…
1 paper · 0 benchmarks
Fraud_Case_Verdicts (The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset)
The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset This data set is based on the judgments of "Offenses of Fraudulence" cases published by the Judicial Yuan.
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
HT Docking is a dataset consisting of 200 million 3D complex structures and 2D structure scores across a consistent set of 13 million in-stock'' molecules over 15 receptors, or binding sites, across the SARS-CoV-2 proteome.
1 paper · 0 benchmarks
HTDM (Hypertention Disease Medication)
Hypertention Disease Medication dataset.
1 paper · 0 benchmarks
IEE is a financial-domain dataset of the Insurance-entity extraction task.
1 paper · 0 benchmarks
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks
JDDC 2.0 is a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform JD.com, containing about 246 thousand dialogue sessions, 3 million utterances, and 507 thousand images, along with…
1 paper · 0 benchmarks
LARQS (An Evaluation Dataset for Chinese Codex Word Embedding Model)
Word embedding is a modern distributed word representations approach widely used in many natural language processing tasks.
1 paper · 0 benchmarks
LARa (Logistic Activity Recognition Challenge)
LARa is the first freely accessible logistics-dataset for human activity recognition.
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
LPBA40 (LONI Probabilistic Brain Atlas)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LSICC (Large Scale Informal Chinese Corpus)
Large Scale Informal Chinese Corpus (LSICC) is a large-scale corpus of informal Chinese.
1 paper · 0 benchmarks
An open-source online generative dictionary that takes a word and context containing the word as input and automatically generates a definition as output.
1 paper · 0 benchmarks
A 160B bilingual long-text dataset with 3 categories: holistic, aggregated and chaotic long texts.
1 paper · 0 benchmarks
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
MMInstruct-GPT4V (MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity)
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: 1.
1 paper · 0 benchmarks
MMTB (Multi-Mission Tool Bench)
Our test data has undergone five rounds of manual inspection and correction by five senior algorithm researcher with years of experience in NLP, CV, and LLM, taking about one month in total.
1 paper · 0 benchmarks
MTC is a financial-domain dataset of the multi-label topic classification task.
1 paper · 0 benchmarks
MUSIED is a large-scale Chinese event detection dataset based on user reviews, text conversations, and phone conversations in a leading e-commerce platform for food service, designed for event detection tasks.
1 paper · 0 benchmarks
MVALUE (Multilingual human VALUE dataset)
Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and…
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.