Datasets › NSC

NSC (National Speech Corpus)

archive 2025-07-28

The National Speech Corpus (NSC) is a significant initiative led by the Info-communications and Media Development Authority (IMDA) of Singapore. It serves as the first large-scale Singapore English corpus and aims to be a valuable resource for researchers and developers working on automatic speech recognition (ASR) technology and other speech-related applications¹²³.

Here are some key points about the National Speech Corpus:

  1. Purpose: The NSC is designed to become an essential source of open speech data for ASR research and speech-related applications.
  2. Importance: By providing a large-scale corpus of Singapore English, it enables advancements in speech recognition, speech synthesis, and other related technologies.
  3. AI and Digital Solutions: The NSC harnesses the power of Artificial Intelligence (AI), paving the way for innovative digital solutions and driving progress in Singapore's digital landscape.
  4. Accent Recognition: Supporting speech technologies can sometimes struggle with recognizing and transcribing locally accented English. The NSC addresses this gap by improving speech engines' accuracy for locally accented English.
  5. Speech Synthesis: The NSC also contributes to speech synthesis technology, allowing AI voices to be produced with more accurate local pronunciations and familiarity for Singaporeans.
  6. Applications: As speech technology improves, ASR can be used in various applications, such as telco call centers, chatbots, and more. For instance:
    • Telco call centers can use ASR to transcribe calls for auditing and sentiment analysis purposes.
    • Chatbots can go beyond text and accurately support our accent while replying in a familiar local accent with accurate pronunciations of street names and food.

The NSC is available under the Singapore Open Data License, and researchers interested in using it can download the corpus via a Dropbox account. The current size of the NSC is approximately 1.2 TB, and it continues to contribute to advancements in speech technology¹³.

Source: Conversation with Bing, 3/17/2024 (1) National Speech Corpus - Infocomm Media Development Authority. https://www.imda.gov.sg/how-we-can-help/national-speech-corpus. (2) ISCA - ArchiveRedirection - isca-speech.org. https://www.isca-speech.org/archive/interspeech_2019/koh19_interspeech.html. (3) National Speech Corpus - Infocomm Media Development Authority. https://www.imda.gov.sg/about-imda/emerging-technologies-and-research/artificial-intelligence/national-speech-corpus.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • NSC

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections