{"url":"/dataset/aidatatang-200zh","name":"aidatatang_200zh","full_name":"aidatatang_200zh","description_markdown":"A Chinese Mandarin speech corpus by Beijing DataTang Technology Co., Ltd, containing 200 hours of speech data from 600 speakers. The transcription accuracy for each sentence is larger than 98%.\r\nAidatatang_200zh is a free Chinese Mandarin speech corpus provided by Beijing DataTang Technology Co., Ltd under Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Public License.\r\nThe contents and the corresponding descriptions of the corpus include:\r\n\r\nThe corpus contains 200 hours of acoustic data, which is mostly mobile recorded data.\r\n600 speakers from different accent areas in China are invited to participate in the recording.\r\nThe transcription accuracy for each sentence is larger than 98%.\r\nRecordings are conducted in a quiet indoor environment.\r\nThe database is divided into training set, validation set, and testing set in a ratio of 7: 1: 2.\r\nDetail information such as speech data coding and speaker information is preserved in the metadata file.\r\nSegmented transcripts are also provided.\r\nThe corpus aims to support researchers in speech recognition, machine translation, voiceprint recognition, and other speech-related fields. Therefore, the corpus is totally free for academic use.","description_withheld":null,"homepage":"http://openslr.org/62/","introduced_date":"2022-06-15","introduced_date_note":null,"introduced_by":null,"license":{"name":"Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)","url":null},"modalities":[{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Speech Recognition","url":"/task/speech-recognition","datasets_with_task":"/datasets/task/speech-recognition"}],"languages":[{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["aidatatang_200zh"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}