{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/english-conversational-telephone-speech","title":"English Conversational Telephone Speech Recognition by Humans and Machines","arxiv_id":"1703.02136","date":"2017-03-06","proceeding":null,"authors":["George Saon","Gakuto Kurata","Tom Sercu","Kartik Audhkhasi","Samuel Thomas","Dimitrios Dimitriadis","Xiaodong Cui","Bhuvana Ramabhadran","Michael Picheny","Lynn-Li Lim","Bergul Roomi","Phil Hall"],"abstract":"One of the most difficult speech recognition tasks is accurate recognition of\nhuman to human communication. Advances in deep learning over the last few years\nhave produced major speech recognition improvements on the representative\nSwitchboard conversational corpus. Word error rates that just a few years ago\nwere 14% have dropped to 8.0%, then 6.6% and most recently 5.8%, and are now\nbelieved to be within striking range of human performance. This then raises two\nissues - what IS human performance, and how far down can we still drive speech\nrecognition error rates? A recent paper by Microsoft suggests that we have\nalready achieved human performance. In trying to verify this statement, we\nperformed an independent set of human performance measurements on two\nconversational tasks and found that human performance may be considerably\nbetter than what was earlier reported, giving the community a significantly\nharder goal to achieve. We also report on our own efforts in this area,\npresenting a set of acoustic and language modeling techniques that lowered the\nword error rate of our own English conversational telephone LVCSR system to the\nlevel of 5.5%/10.3% on the Switchboard/CallHome subsets of the Hub5 2000\nevaluation, which - at least at the writing of this paper - is a new\nperformance milestone (albeit not at what we measure to be human performance!).\nOn the acoustic side, we use a score fusion of three models: one LSTM with\nmultiple feature inputs, a second LSTM trained with speaker-adversarial\nmulti-task learning and a third residual net (ResNet) with 25 convolutional\nlayers and time-dilated convolutions. On the language modeling side, we use\nword and character LSTMs and convolutional WaveNet-style language models.","url_abs":"http://arxiv.org/abs/1703.02136v1","url_pdf":"http://arxiv.org/pdf/1703.02136v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-recognition-on-switchboard-hub500","task":"Speech Recognition","dataset":"Switchboard + Hub500","model":"ResNet + BiLSTMs acoustic model","rank_in_archive_order":3,"of":30,"metrics":{"Percentage error":"5.5"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-swb_hub_500-wer","task":"Speech Recognition","dataset":"swb_hub_500 WER fullSWBCH","model":"ResNet + BiLSTMs acoustic model","rank_in_archive_order":3,"of":12,"metrics":{"Percentage error":"10.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.02136","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}