Home › Datasets › language › Chinese

Chinese datasets

archive 2025-07-28

460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 10 of 10: 28 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Chinese datasets 433–460 of 460

This dataset is used for user identity linkage across two online social networks in Chinese.
1 paper · 0 benchmarks
Wiki-zh is an annotated Chinese dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
XiaChuFang Recipe Corpus contains recipes are from 下厨房 (XiaChuFang), a popular Chinese recipe sharing website.
1 paper · 0 benchmarks
IMO-level geometry problem with complete natural language description, geometric shapes, formal language annotations, and theorem sequences annotations.
1 paper · 0 benchmarks
fruit-SALAD is a synthetic image dataset with 10,000 generated images of fruit depictions.
1 paper · 0 benchmarks
titanic5 Dataset Created by David Beltran del Rio March 2016.
1 paper · 0 benchmarks
AntM2C (Ant-Group Multi-Scenario Multi-Modal CTR dataset)
We release a large-scale Multi-Scenario Multi-Modal CTR dataset named AntM2C, built from real industrial data from Alipay.
0 papers · 0 benchmarks
CBLPRD-330k (China-Balanced-License-Plate-Recognition-Dataset-330k)
A high-quality, balanced dataset of 330,000 images featuring various types of Chinese license plates.
0 papers · 0 benchmarks
We introduce CCI4.0, a large-scale bilingual pre-training dataset engineered for superior data quality and diverse human-like reasoning trajectory.
0 papers · 0 benchmarks
CNTD (Chinese and Naxi text detection)
Chinese and Naxi scene text detection data set, labelme to json.
0 papers · 0 benchmarks
Feifei (FeifeiY)
potential use
0 papers · 0 benchmarks
A high-quality dataset forms the foundation for machine learning-based predictions of structural load capacity.
0 papers · 0 benchmarks
GDXray+ is a collection of more than 21.100 X-ray images for the development, testing, and evaluation of image analysis and computer vision algorithms.
0 papers · 0 benchmarks
The Hong Kong Cantonese Corpus was collected from transcribed conversations that were recorded between March 1997 and August 1998.
0 papers · 0 benchmarks
LSARS (Large Scale Abstractive multi-Review Summarization)
In an active e-commerce environment, customers process a large number of reviews when deciding on whether to buy a product or not.
0 papers · 0 benchmarks
IQ testing has served as a foundational methodology for evaluating human cognitive capabilities, deliberately decoupling assessment from linguistic background, language proficiency, or domain-specific knowledge to isolate core competencies…
0 papers · 0 benchmarks
OpenSurfaces is a large database of annotated surfaces created from real-world consumer photographs.
0 papers · 0 benchmarks
PCN (Pedestrian Color Naming)
Pedestrian Color Naming (PCN) is a dataset for pedestrian color naming, which contains 14,213 images, each of which hand-labeled with color label for each pixel.
0 papers · 0 benchmarks
Shanghai2020 (Shanghai-2020 Dataset)
It is released by the Shanghai Central Meteorological Observatory (SCMO) in 2020, records serval years of historical precipitation events in the Yangtze River delta area.
0 papers · 0 benchmarks
SportsSum is a Chinese sports game summarization dataset that contains 5,428 soccer games of live commentaries and the corresponding news articles.
0 papers · 0 benchmarks
State Farm (State Farm Distracted Driver Detection)
该数据集是一个全面而多样化的驾驶员行为监测数据集,其中包括来自美洲,亚洲和非洲的26名不同种族,肤色和性别(13名男性和13名女性)的参与者。数据集中的所有图像都是由固定在汽车仪表板上的摄像头拍摄的,所有图像都是RGB像素。该数据集共包含22424张图像。
0 papers · 0 benchmarks
TCMP-300 (Traditional Chinese Medicinal Plant Dataset)
Traditional Chinese medicinal plants are often used to prevent and treat diseases for the human body.
0 papers · 1 benchmark
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
WalnutData (A UAV Remote Sensing Dataset of Green Walnuts and Model Evaluation)
With the gradual maturity of UAV technology, it can provide extremely powerful support for smart agriculture and precise monitoring.
0 papers · 0 benchmarks
A Chinese Mandarin speech corpus by Beijing DataTang Technology Co., Ltd, containing 200 hours of speech data from 600 speakers.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.