Home › Datasets › language › Chinese

Chinese datasets

archive 2025-07-28

460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 6 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Chinese datasets 241–288 of 460

DiaASQ (Conversational Aspect-based Sentiment Quadruple Extraction)
DiaASQ is a fine-grained Aspect-based Sentiment Analysis (ABSA) benchmark under the conversation scenario.
4 papers · 2 benchmarks
Diamante is a novel and efficient framework consisting of a data collection strategy and a learning method to boost the performance of pre-trained dialogue models.
4 papers · 0 benchmarks
FRMT (Few-shot Region-aware Machine Translation)
FRMT is a dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation.
4 papers · 4 benchmarks
Pretrain: 200k Instruction: 100k
4 papers · 0 benchmarks
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
M³-VOS (M³-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation)
💡 Description A new benchmark, Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation (M³-VOS), to verify the ability of models to understand object phases, which consists of 479 high-resolution videos spanning over 10…
4 papers · 1 benchmark
MultiSpider is a large multilingual text-to-SQL dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese).
4 papers · 0 benchmarks
SLING (Sino LINGuistics)
SLING consists of 38K minimal sentence pairs in Mandarin Chinese grouped into 9 high-level linguistic phenomena.
4 papers · 0 benchmarks
The WikiSem500 dataset contains around 500 per-language cluster groups for English, Spanish, German, Chinese, and Japanese (a total of 13,314 test cases).
4 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
Youku-mPLUG is a large Chinese high-quality video-language dataset which is collected from Youku.com, a well-known Chinese video-sharing website, with strict criteria of safety, diversity, and quality.
4 papers · 0 benchmarks
The Aircraft Context Dataset, a composition of two inter-compatible large-scale and versatile image datasets focusing on manned aircraft and UAVs, is intended for training and evaluating classification, detection and segmentation models in…
3 papers · 0 benchmarks
BenchIE: a benchmark and evaluation framework for comprehensive evaluation of OIE systems for English, Chinese and German.
3 papers · 1 benchmark
CNewSum is a large-scale Chinese news summarization dataset which consists of 304,307 documents and human-written summaries for the news feed.
3 papers · 0 benchmarks
A benchmark dataset with 960 pairs of Chinese wOrd Similarity, where all the words have two morphemes in three Part of Speech (POS) tags with their human annotated similarity rather than relatedness.
3 papers · 0 benchmarks
CPP (Chinese Polyphones with Pinyin)
A benchmark dataset that consists of 99,000+ sentences for Chinese polyphone disambiguation.
3 papers · 1 benchmark
The high-quality multi-turn dialogue dataset, which has a total of 3,134 multi-turn consultation dialogues.
3 papers · 0 benchmarks
CS (Chinese Simile)
This dataset is constructed and based on the online free-access fictions that are tagged with sci-fi, urban novel, love story, youth, etc.
3 papers · 0 benchmarks
CV-Cities comprises $223,736$ ground panoramic images and an equal number of satellite images all accompanied by high-precision GPS coordinates.
3 papers · 1 benchmark
DISRPT2021 (DISRPT2021 shared task on Discourse Unit Segmentation, Connective Detection and Discourse Relation Classification)
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism…
3 papers · 0 benchmarks
This is a medical multiple-choice dataset with explanations which can be used to interpret the answer.
3 papers · 0 benchmarks
HUME-VB (The Hume Vocal Bursts Dataset)
The Hume Vocal Burst Database (H-VB) includes all train, validation, and test recordings and corresponding emotion ratings for the train and validation recordings.
3 papers · 7 benchmarks
ImageNet_CN (Chinese ImageNet Classification)
transform the ImageNet-1K classification datatset for Chinese models by translating labels and prompts into Chinese.
3 papers · 1 benchmark
Collected by cleaning data from knowledge-intensive websites like Wikipedia and science and technology reports, and processing it using reverse engineering techniques.
3 papers · 0 benchmarks
MISP2021 (Multimodal Information Based Speech Processing 2021)
The MISP2021 challenge dataset is a collection of audio-visual conversational data recorded in a home TV scenario using distant multi-microphones.
3 papers · 0 benchmarks
MULTI-Benchmark is a cutting-edge benchmark for evaluating Multimodal Large Language Models (MLLMs).
3 papers · 0 benchmarks
ODSQA (Open-Domain Spoken Question Answering)
The ODSQA dataset is a spoken dataset for question answering in Chinese.
3 papers · 0 benchmarks
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
3 papers · 1 benchmark
TT100K (Tsinghua-Tencent 100K(official training and testing set))
Trainging and testing data: The original training set includes 6105 images, and the original testing set includes 3071 images.
3 papers · 1 benchmark
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
Title2Event is a large-scale sentence-level dataset for benchmarking Open Event Extraction without restricting event types.
3 papers · 0 benchmarks
WDC-Dialogue is a dataset built from the Chinese social media to train EVA.
3 papers · 0 benchmarks
Dataset Description The dataset described in the provided text is focused on social media polls collected from Weibo, a popular Chinese microblogging platform.
3 papers · 3 benchmarks
Wikipedia Title is a dataset for learning character-level compositionality from the character visual characteristics.
3 papers · 0 benchmarks
6981 SAT-level geometry problem with complete natural language description, geometric shapes, formal language annotations, and theorem sequences annotations.
3 papers · 0 benchmarks
Our trajectory dataset consists of camera-based images, LiDAR scanned point clouds, and manually annotated trajectories.
2 papers · 1 benchmark
CA4P-483 is a dataset designed to facilitate the sequence labeling tasks and regulation compliance identification between privacy policies and software.
2 papers · 0 benchmarks
CII-Bench (Chinese Image Implication understanding Benchmark)
We introduce the Chinese Image Implication Understanding Benchmark CII-Bench, a new benchmark measuring the higher-order perceptual, reasoning and comprehension abilities of MLLMs when presented with complex Chinese implication images.
2 papers · 0 benchmarks
The general multi-turn dialogue evaluation dataset with nine topics.
2 papers · 0 benchmarks
Chinese Spelling Correction Dataset for errors generated by pinyin IME (CSCD-IME), a dataset containing 40,000 annotated sentences from real posts of official media on Sina Weibo.
2 papers · 0 benchmarks
We introduce ChinaTravel, the first open-ended benchmark grounded in authentic Chinese travel requirements collected from 1,154 human participants.
2 papers · 0 benchmarks
Classifiers are function words that are used to express quantities in Chinese and are especially difficult for language learners.
2 papers · 0 benchmarks
Chinese Gigaword corpus consists of 2.2M of headline-document pairs of news stories covering over 284 months from two Chinese newspapers, namely the Xinhua News Agency of China (XIN) and the Central News Agency of Taiwan (CNA).
2 papers · 0 benchmarks
DISC-Law-SFT comprises two subsets, DISC-Law-SFT-Pair and DISC-Law-SFT-Triplet.
2 papers · 0 benchmarks
DialogUSR dataset covers 23 domains with a multi-step crowd-sourcing procedure.
2 papers · 0 benchmarks
ExpMRC is a benchmark for the Explainability evaluation of Machine Reading Comprehension.
2 papers · 0 benchmarks
Fashion-MNT is large-scale bilingual product description dataset called Fashion-MMT, which contains over 114k noisy and 40k manually cleaned description translations with multiple product images.
2 papers · 0 benchmarks
GBUSV (Gallbladder Ultrasound Videos)
Description GBUSV is a un-annotated dataset consisting of ultrasound videos of of patients with either of a malignant or a non-malignant gallbladder.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.