{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/realizing-petabyte-scale-acoustic-modeling","title":"Realizing Petabyte Scale Acoustic Modeling","arxiv_id":"1904.10584","date":"2019-04-24","proceeding":null,"authors":["Sree Hari Krishnan Parthasarathi","Nitin Sivakrishnan","Pranav Ladkat","Nikko Strom"],"abstract":"Large scale machine learning (ML) systems such as the Alexa automatic speech\nrecognition (ASR) system continue to improve with increasing amounts of\nmanually transcribed training data. Instead of scaling manual transcription to\nimpractical levels, we utilize semi-supervised learning (SSL) to learn acoustic\nmodels (AM) from the vast firehose of untranscribed audio data. Learning an AM\nfrom 1 Million hours of audio presents unique ML and system design challenges.\nWe present the design and evaluation of a highly scalable and resource\nefficient SSL system for AM. Employing the student/teacher learning paradigm,\nwe focus on the student learning subsystem: a scalable and robust data pipeline\nthat generates features and targets from raw audio, and an efficient model\npipeline, including the distributed trainer, that builds a student model. Our\nevaluations show that, even without extensive hyper-parameter tuning, we obtain\nrelative accuracy improvements in the 10 to 20$\\%$ range, with higher gains in\nnoisier conditions. The end-to-end processing time of this SSL system was 12\ndays, and several components in this system can trivially scale linearly with\nmore compute resources.","url_abs":"http://arxiv.org/abs/1904.10584v1","url_pdf":"http://arxiv.org/pdf/1904.10584v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"realizing-petabyte-scale-acoustic-modeling","repo_url":"https://github.com/OpenSourceAI/sota-server","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"am","method_name":"AM"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}