{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spgispeech-5000-hours-of-transcribed","title":"SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition","arxiv_id":"2104.02014","date":"2021-04-05","proceeding":null,"authors":["Patrick K. O'Neill","Vitaly Lavrukhin","Somshubra Majumdar","Vahid Noroozi","Yuekai Zhang","Oleksii Kuchaiev","Jagadeesh Balam","Yuliya Dovzhenko","Keenan Freyberg","Michael D. Shulman","Boris Ginsburg","Shinji Watanabe","Georg Kucsko"],"abstract":"In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitalization, punctuation, and denormalization of non-standard words) is imputed by separate post-processing models. This adds complexity and limits performance, as many formatting tasks benefit from semantic information present in the acoustic signal but absent in transcription. Here we propose a new STT task: end-to-end neural transcription with fully formatted text for target labels. We present baseline Conformer-based models trained on a corpus of 5,000 hours of professionally transcribed earnings calls, achieving a CER of 1.7. As a contribution to the STT research community, we release the corpus free for non-commercial use at https://datasets.kensho.com/datasets/scribe.","url_abs":"https://arxiv.org/abs/2104.02014v2","url_pdf":"https://arxiv.org/pdf/2104.02014v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spgispeech-5000-hours-of-transcribed","repo_url":"https://github.com/espnet/espnet/tree/master/egs2/spgispeech/asr1","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-to-text","task_name":"Speech-to-Text"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[{"slug":"spgispeech","name":"SPGISpeech","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-recognition-on-spgispeech","task":"Speech Recognition","dataset":"SPGISpeech","model":"Conformer","rank_in_archive_order":3,"of":3,"metrics":{"Word Error Rate (WER)":"5.7"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2104.02014","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}