Home › Datasets › language › Mandarin Chinese

Mandarin Chinese datasets

archive 2025-07-28

25 datasets carry the language tag "Mandarin Chinese", ordered by the archive's paper count. Page 1 of 1: 25 shown of 25. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Mandarin Chinese datasets 1–25 of 25

Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
AISHELL-1 is a corpus for speech recognition research and building speech recognition systems for Mandarin.
197 papers · 1 benchmark
AISHELL-2 contains 1000 hours of clean read-speech data from iOS is free for academic usage.
75 papers · 4 benchmarks
ACE 2005 (ACE 2005 Multilingual Training Corpus)
ACE 2005 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2005 Automatic Content Extraction (ACE) technology evaluation.
65 papers · 8 benchmarks
WenetSpeech is a multi-domain Mandarin corpus consisting of 10,000+ hours high-quality labeled speech, 2,400+ hours weakly labelled speech, and about 10,000 hours unlabeled speech, with 22,400+ hours in total.
58 papers · 1 benchmark
ACE 2004 (ACE 2004 Multilingual Training Corpus)
ACE 2004 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2004 Automatic Content Extraction (ACE) technology evaluation.
51 papers · 6 benchmarks
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems.
38 papers · 0 benchmarks
The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.
12 papers · 0 benchmarks
BiPaR is a manually annotated bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support monolingual, multilingual and cross-lingual reading comprehension on novels.
6 papers · 0 benchmarks
FMFCC-A is a large publicly-available Mandarin dataset for synthetic speech detection, which contains 40,000 synthesized Mandarin utterances that generated by 11 Mandarin TTS systems and two Mandarin VC systems, and 10,000 genuine Mandarin…
6 papers · 0 benchmarks
CSRC (Children Speech Recognition Challenge)
CSRC is a collection of data for Children Speech Recognition.
5 papers · 0 benchmarks
DISRPT2019 (DISRPT2019 shared task on Discourse Unit Segmentation and Connective Detection)
The DISRPT 2019 workshop introduces the first iteration of a cross-formalism shared task on discourse unit segmentation.
4 papers · 0 benchmarks
EMOVIE is a Mandarin emotion speech dataset including 9,724 samples with audio files and its emotion human-labeled annotation.
3 papers · 0 benchmarks
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants.
3 papers · 0 benchmarks
xMIND (A Multilingual Dataset for Cross-lingual News Recommendation)
xMIND is an open, large-scale multilingual news dataset for multi- and cross-lingual news recommendation.
3 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
K-SportsSum is a sports game summarization dataset with two characteristics: (1) K-SportsSum collects a large amount of data from massive games.
2 papers · 0 benchmarks
CCSE (Chinese Character Stroke Extraction)
Chinese Character Stroke Extraction (CCSE) is a benchmark containing two large-scale datasets: Kaiti CCSE (CCSE-Kai) and Handwritten CCSE (CCSE-HW).
1 paper · 0 benchmarks
NL2GQL Dataset (NL2GQL developped for R3-NL2GQL)
A bilingual (English and Chinese natural language queries) dataset which has NL queries annotated with their corresponding GQL queries (i.e.
1 paper · 0 benchmarks
This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh).
1 paper · 0 benchmarks
Engagement with the government of Taiwan as part of the vTaiwan participatory process which led to the successful regulation of Uber in Taiwan.
1 paper · 0 benchmarks
Poly-FEVER is a multilingual fact verification benchmark designed to evaluate hallucination detection in large language models (LLMs).
1 paper · 0 benchmarks
UNER v1 (Universal NER v1)
UNER v1 adds an NER annotation layer to 18 datasets (primarily treebanks from UD) and covers 12 geneologically and ty- pologically diverse languages: Cebuano, Danish, German, English, Croatian, Portuguese, Russian, Slovak, Serbian,…
1 paper · 31 benchmarks
mDRT (Multilingual Diagnostic Rhyme Test)
We present a multilingual test set for conducting speech intelligibility tests in the form of diagnostic rhyme tests.
1 paper · 0 benchmarks
MCCSD (Mandarin Chinese Cued Speech Dataset)
This MCCS dataset is the first large-scale Mandarin Chinese Cued Speech dataset.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.