Home › Datasets › language › Chinese

Chinese datasets

archive 2025-07-28

460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 9 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Chinese datasets 385–432 of 460

MVME (Multi-View Medical Evaluation Benchmark)
The benchmark assesses the real-time interactive consultation capabilities of LLMs across three critical dimensions.
1 paper · 0 benchmarks
Dataset introduction There are four dimension in MBTI.
1 paper · 0 benchmarks
Existing arithmetic benchmarks have a limited number of multiple-choice questions.
1 paper · 1 benchmark
Existing arithmetic benchmarks have a limited number of True-or-False questions.
1 paper · 1 benchmark
Collecting data with a HIKVISION USB Camera DS-E11, we build a dataset called MentalHAD with four abnormal actions (climbing walls, hitting windows, climbing, and hitting) and six normal actions (crouching, standing, sitting, hand waving,…
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
https://huggingface.co/datasets/OpenDFM/MobA-MobBench
1 paper · 0 benchmarks
M²ConceptBase is a concept-centric multimodal knowledge base designed to bridge the gap between visual and linguistic semantics.
1 paper · 0 benchmarks
NAIST COVID is a multilingual dataset of social media posts related to COVID-19, consisting of microblogs in English and Japanese from Twitter and those in Chinese from Weibo.
1 paper · 0 benchmarks
NBMOD (Noisy Background Multi-Object Dataset for grasp detection)
Introduction NBMOD is a dataset created for researching the task of specific object grasp detection by robots in noisy environments.
1 paper · 1 benchmark
NPO (Negative and Positive Obstacles)
The dataset is recorded with an on-vehicle ZED stereo camera in both urban and rural environments The dataset contains various lighting conditions, such as normal lights, large-area shadows, dim lights, and sun glare.
1 paper · 1 benchmark
NaSGEC is a new dataset to facilitate research on Chinese grammatical error correction (CGEC) for native speaker texts from multiple domains.
1 paper · 0 benchmarks
Introduction to the Needle In A Haystack Test The Needle In A Haystack test, inspired by NeedleInAHaystack, is an evaluation method that randomly inserts key information into long texts to create prompts for large language models (LLMs).
1 paper · 0 benchmarks
Collected by cleaning data from daily Xinwen Lianbo transcripts over the past three months and processing it using reverse engineering techniques.
1 paper · 0 benchmarks
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
1 paper · 0 benchmarks
Dataset OQRanD and OQGenD for paper "Asking the crowd: Asking the Crowd: Question Analysis, Evaluation and Generation for Open Discussion on Online Forums" by Zi Chai, Xinyu Xing, Xiaojun Wan and Bo Huang.
1 paper · 0 benchmarks
Dataset OQRanD and OQGenD for paper "Asking the crowd: Asking the Crowd: Question Analysis, Evaluation and Generation for Open Discussion on Online Forums" by Zi Chai, Xinyu Xing, Xiaojun Wan and Bo Huang.
1 paper · 0 benchmarks
OSIC Pulmonary Fibrosis Progression (OSIC Pulmonary Fibrosis Progression: Predict lung function decline)
Imagine one day, your breathing became consistently labored and shallow.
1 paper · 0 benchmarks
Orchid2024 is a fine-grained classification dataset specifically designed for Chinese Cymbidium orchid cultivars.
1 paper · 0 benchmarks
PHP Webshell Dataset (First PHP Webshell Opcode Incremental Dataset)
First PHP Webshell Opcode Incremental Dataset Motivation To improve the robustness of PHP webshell detection by analyzing low-level opcode patterns, circumventing common code obfuscation and evasion techniques.
1 paper · 0 benchmarks
PSM is a financial-domain dataset of the pairwise search matching task.
1 paper · 0 benchmarks
PTVD is a plot-oriented multimodal dataset in the TV domain.
1 paper · 0 benchmarks
Pan+ChiPhoto dataset is a Chinese character dataset.
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
A high-quality large-scale dataset consisting of 49,000+ data samples for the task of Chinese query-based document summarization.
1 paper · 0 benchmarks
Multilingual explainable fact-checking dataset on Russia-Ukraine Conflict 2022
1 paper · 0 benchmarks
SAS-Bench represents the first specialized benchmark for evaluating Large Language Models (LLMs) on Short Answer Scoring (SAS) tasks.
1 paper · 0 benchmarks
A Chinese sign language dataset that includes dialogue information.
1 paper · 0 benchmarks
SSD_ID (Sub-Slot Dialogue dataset id number domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 0 benchmarks
SSD_NAME (Sub-Slot Dialogue dataset name domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 1 benchmark
SSD_PLATE (Sub-Slot Dialogue dataset license plate number domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 0 benchmarks
With the rise of social media, user-generated content has surged, and hate speech has proliferated.
1 paper · 0 benchmarks
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts 🔥 Key Features - 3000+ hours of synthetic speech - Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including: - 📖 Reading Style - 🎙️…
1 paper · 0 benchmarks
Photometric stereoscopic test data sets under six lights taken using laboratory equipment.
1 paper · 0 benchmarks
The datasets used in our WACV paper High-Fidelity Document Stain Removal via A Large-Scale Real-World Dataset and A Memory-Augmented Transformer.
1 paper · 0 benchmarks
Steel Tube Dataset (Steel Tube Weld Defect Detection Dataset)
8 kinds of weld defects
1 paper · 0 benchmarks
TQBA++ (Tiny QA Benchmark++)
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
Please refer this paper @article{sarker2024tea, author = {Sarker, Swapnil Sharma and Islam, Ashiqul and Talukder Raktim, Raufun and Roshni, Sanjana and Joy, Sajib Kumar Saha and Shah, Faisal}, title = {Real-Time Tea Leaf Disease Detection…
1 paper · 0 benchmarks
This dataset consists of 2,192 high-quality traditional Chinese landscape paintings (中国山水画).
1 paper · 0 benchmarks
https://github.com/zzr-idam/Under-Display-Camera-UAV
1 paper · 0 benchmarks
UHGEvalDataset contains over 5000 news items.
1 paper · 0 benchmarks
UNER v1 (Universal NER v1)
UNER v1 adds an NER annotation layer to 18 datasets (primarily treebanks from UD) and covers 12 geneologically and ty- pologically diverse languages: Cebuano, Danish, German, English, Croatian, Portuguese, Russian, Slovak, Serbian,…
1 paper · 31 benchmarks
This task stems from the observation that text embedded in images is intrinsically different from common visual elements and natural language due to the need to align the modalities of vision, text, and text embedded in images.
1 paper · 0 benchmarks
A high-resolution version of VGGFace2 for academic face editing purposes.
1 paper · 0 benchmarks
VTQA (Visual Text Question Answering)
VTQA is a dataset containing open-ended questions about image-text pairs.
1 paper · 0 benchmarks
Voice Navigation is a large-scale dataset of Chinese speech for slot filling, containing more than 830,000 samples.
1 paper · 0 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.