Home › Datasets › language › Chinese

Chinese datasets

archive 2025-07-28

460 datasets carry the language tag "Chinese", ordered by the archive's paper count. Page 7 of 10: 48 shown of 460. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Chinese datasets 289–336 of 460

Hansel is a human-annotated Chinese entity linking (EL) dataset, focusing on tail entities and emerging entities: - The test set contains Few-shot (FS) and zero-shot (ZS) slices, has 10K examples and uses Wikidata as the corresponding…
2 papers · 0 benchmarks
InfiniteBench (∞Bench: Extending Long Context Evaluation Beyond 100K Tokens)
Introduction Welcome to InfiniteBench, a cutting-edge benchmark tailored for evaluating the capabilities of language models to process, understand, and reason over super long contexts (100k+ tokens).
2 papers · 0 benchmarks
K-SportsSum is a sports game summarization dataset with two characteristics: (1) K-SportsSum collects a large amount of data from massive games.
2 papers · 0 benchmarks
The Live Comment Dataset is a large-scale dataset with 2,361 videos and 895,929 live comments that were written while the videos were streamed.
2 papers · 0 benchmarks
M2QA (Multi-domain Multilingual Question Answering)
M2QA (Multi-domain Multilingual Question Answering) is an extractive question answering benchmark for evaluating joint language and domain transfer.
2 papers · 0 benchmarks
MCSCSet is a large-scale specialist-annotated dataset, designed for the task of Medical-domain Chinese Spelling Correction that contains about 200k samples.
2 papers · 0 benchmarks
MGSM8KInstruct, the multilingual math reasoning instruction dataset, encompassing ten distinct languages, thus addressing the issue of training data scarcity in multilingual math reasoning.
2 papers · 0 benchmarks
MSDA (Multi-source domain adaptation dataset for text recognition)
5 domains: synthetic domain, document domain, street view domain, handwritten domain, and car license domain over five million images
2 papers · 2 benchmarks
MiniWob++ is a suite of web-browser based tasks introduced in Liu et al.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
NoW (Noise of Web)
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
2 papers · 0 benchmarks
OIR is a financial-domain dataset of the outbound intent recognition task.
2 papers · 0 benchmarks
PETCI (PETCI: A Parallel English Translation Dataset of Chinese Idioms)
PETCI is a Parallel English Translation dataset of Chinese Idioms, collected from an idiom dictionary and Google and DeepL translation.
2 papers · 0 benchmarks
The PKU dataset has almost 4,000 images categorized into five groups (G1-G5) that show different situations.
2 papers · 0 benchmarks
PNPB dataset (Pig Novelty Preference Behavior dataset)
The dataset consists of a total of 20 videos, each of which is 5.5 minutes long in duration.
2 papers · 0 benchmarks
The Parallel Meaning Bank (PMB), developed at the University of Groningen and building upon the Groningen Meaning Bank, comprises sentences and texts in raw and tokenised format, syntactic analysis, word senses, thematic roles, reference…
2 papers · 0 benchmarks
Perseus is a dataset for Cross-Lingual Summarization (CLS) which collects about 94K Chinese scientific documents paired with English summaries.
2 papers · 0 benchmarks
ProSLU (Profile-based Spoken Language Understanding)
In the paper, to bridge the research gap, we propose a new and important task, Profile-based Spoken Language Understanding (ProSLU), which requires a model not only depends on the text but also on the given supporting profile information.
2 papers · 2 benchmarks
A human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
2 papers · 0 benchmarks
SSD (Sub-Slot Dialogue dataset)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
2 papers · 0 benchmarks
SSD_PHONE (Sub-Slot Dialogue dataset phone domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
2 papers · 0 benchmarks
SVBench (Streaming Video Understanding Benchmark)
Dataset Card for SVBench This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources.
2 papers · 0 benchmarks
Sakuga-42M is a large-scale hand-drawn cartoon video dataset for academic research purposes, it comprises 42 million cartoon keyframes covering various artistic styles, regions, and years, with comprehensive semantic annotations including…
2 papers · 0 benchmarks
A newly developed natural scene text dataset of Chinese shop signs in street views.
2 papers · 0 benchmarks
SpCQL (Text-to-CQL)
The first dataset contains annotated natural language queries (i.e.
2 papers · 0 benchmarks
A new text effects dataset with 141,081 text effect/glyph pairs in total.
2 papers · 0 benchmarks
VGaokao is a verification style reading comprehension dataset designed for native speakers' evaluation.
2 papers · 0 benchmarks
WeatherKITTI is currently the most realistic all-weather simulated enhancement of the KITTI dataset.
2 papers · 0 benchmarks
Weibo-COV is a large-scale COVID-19 social media dataset from Weibo, covering more than 30 million posts from 1 November 2019 to 30 April 2020.
2 papers · 0 benchmarks
XL-R2R (Cross-lingual Room-to-Room)
The XL-R2R dataset is built upon the R2R dataset and extends it with Chinese instructions.
2 papers · 0 benchmarks
AODRaw (Adverse condition Object Detection with RAW images)
We introduce the AODRaw dataset, which offers 7,785 high-resolution real RAW images with 135,601 annotated instances spanning 62 categories, capturing a broad range of indoor and outdoor scenes under 9 distinct light and weather conditions.
1 paper · 1 benchmark
AQL-22 (Archive Query Log)
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
A Rich Annotated Mandarin Conversational (RAMC) Speech Dataset, including 180 hours of Mandarin Chinese dialogue, 150, 10 and 20 hours for the training set, development set and test set respectively.
1 paper · 0 benchmarks
ASRD (Anime Style Recognition Dataset)
A well-labeled challenging dataset, to facilitate the research on style recognition on anime images by collecting images from 190 anime and cartoon works covering 93 years from 13 countries and regions, 2D and 3D work into consideration…
1 paper · 0 benchmarks
The existing multi-modality image fusion dataset lacks comprehensive coverage of adverse weather scenarios.
1 paper · 0 benchmarks
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
1 paper · 0 benchmarks
The Inpainting dataset consists of synchronized Labeled image and LiDAR scanned point clouds.
1 paper · 1 benchmark
The AppealCase dataset is the first large-scale resource specifically designed to support LegalAI research in appellate judgment scenarios.
1 paper · 0 benchmarks
The BWB corpus consists of Chinese novels translated by experts into English, and the annotated test set is designed to probe the ability of machine translation systems to model various discourse phenomena.
1 paper · 0 benchmarks
Baidu PersonaChat, which is a personalization dataset collected and open-sourced by Baidu, is similar to ConvAI2, although it’s Chinese.
1 paper · 0 benchmarks
Overview This is a dataset of blood cells photos.
1 paper · 0 benchmarks
CAGUI (Chinese Android GUI Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data.
1 paper · 3 benchmarks
CC-Riddle (Chinese Character Riddle)
CC-Riddle is a Chinese character riddle dataset covering the majority of common simplified Chinese characters by crawling riddles from the Web and generating brand new ones.
1 paper · 0 benchmarks
To address the scarcity of high-quality safety datasets in the Chinese, we open-sourced the CCI (Chinese Corpora Internet) dataset on November 29, 2023.
1 paper · 0 benchmarks
CCSE (Chinese Character Stroke Extraction)
Chinese Character Stroke Extraction (CCSE) is a benchmark containing two large-scale datasets: Kaiti CCSE (CCSE-Kai) and Handwritten CCSE (CCSE-HW).
1 paper · 0 benchmarks
- CFEVER is a Chinese Fact Extraction and VERification dataset published at AAAI 2024.
1 paper · 0 benchmarks
CHIP Clinical Trial Classification, a dataset aimed at classifying clinical trials eligibility criteria, which are fundamental guidelines of clinical trials defined to identify whether a subject meets a clinical trial or not, is used for…
1 paper · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.