{"url":"/dataset/nsc","name":"NSC","full_name":"National Speech Corpus","description_markdown":"The **National Speech Corpus (NSC)** is a significant initiative led by the **Info-communications and Media Development Authority (IMDA)** of Singapore. It serves as the **first large-scale Singapore English corpus** and aims to be a valuable resource for researchers and developers working on **automatic speech recognition (ASR)** technology and other speech-related applications¹²³.\r\n\r\nHere are some key points about the **National Speech Corpus**:\r\n\r\n1. **Purpose**: The NSC is designed to become an essential source of **open speech data** for ASR research and speech-related applications.\r\n2. **Importance**: By providing a large-scale corpus of Singapore English, it enables advancements in **speech recognition**, **speech synthesis**, and other related technologies.\r\n3. **AI and Digital Solutions**: The NSC harnesses the power of **Artificial Intelligence (AI)**, paving the way for innovative digital solutions and driving progress in Singapore's digital landscape.\r\n4. **Accent Recognition**: Supporting speech technologies can sometimes struggle with recognizing and transcribing **locally accented English**. The NSC addresses this gap by improving speech engines' accuracy for locally accented English.\r\n5. **Speech Synthesis**: The NSC also contributes to **speech synthesis technology**, allowing AI voices to be produced with more accurate local pronunciations and familiarity for Singaporeans.\r\n6. **Applications**: As speech technology improves, ASR can be used in various applications, such as **telco call centers**, **chatbots**, and more. For instance:\r\n    - Telco call centers can use ASR to transcribe calls for **auditing** and **sentiment analysis** purposes.\r\n    - Chatbots can go beyond text and accurately support our accent while replying in a familiar local accent with accurate pronunciations of street names and food.\r\n\r\nThe NSC is available under the **Singapore Open Data License**, and researchers interested in using it can download the corpus via a **Dropbox account**. The current size of the NSC is approximately **1.2 TB**, and it continues to contribute to advancements in speech technology¹³.\r\n\r\nSource: Conversation with Bing, 3/17/2024\r\n(1) National Speech Corpus - Infocomm Media Development Authority. https://www.imda.gov.sg/how-we-can-help/national-speech-corpus.\r\n(2) ISCA - ArchiveRedirection - isca-speech.org. https://www.isca-speech.org/archive/interspeech_2019/koh19_interspeech.html.\r\n(3) National Speech Corpus - Infocomm Media Development Authority. https://www.imda.gov.sg/about-imda/emerging-technologies-and-research/artificial-intelligence/national-speech-corpus.","description_withheld":null,"homepage":"https://www.imda.gov.sg/about-imda/emerging-technologies-and-research/artificial-intelligence/national-speech-corpus","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["NSC"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}