Datasets › aidatatang_200zh

aidatatang_200zh

15 Jun 2022 archive 2025-07-28

A Chinese Mandarin speech corpus by Beijing DataTang Technology Co., Ltd, containing 200 hours of speech data from 600 speakers. The transcription accuracy for each sentence is larger than 98%. Aidatatang_200zh is a free Chinese Mandarin speech corpus provided by Beijing DataTang Technology Co., Ltd under Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Public License. The contents and the corresponding descriptions of the corpus include:

The corpus contains 200 hours of acoustic data, which is mostly mobile recorded data. 600 speakers from different accent areas in China are invited to participate in the recording. The transcription accuracy for each sentence is larger than 98%. Recordings are conducted in a quiet indoor environment. The database is divided into training set, validation set, and testing set in a ratio of 7: 1: 2. Detail information such as speech data coding and speaker information is preserved in the metadata file. Segmented transcripts are also provided. The corpus aims to support researchers in speech recognition, machine translation, voiceprint recognition, and other speech-related fields. Therefore, the corpus is totally free for academic use.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • aidatatang_200zh

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections