{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-speech-recognition-with-adaptive","title":"End-to-end Speech Recognition with Adaptive Computation Steps","arxiv_id":"1808.10088","date":"2018-08-30","proceeding":null,"authors":["Mohan Li","Min Liu","Masanori Hattori"],"abstract":"In this paper, we present Adaptive Computation Steps (ACS) algo-rithm, which\nenables end-to-end speech recognition models to dy-namically decide how many\nframes should be processed to predict a linguistic output. The model that\napplies ACS algorithm follows the encoder-decoder framework, while unlike the\nattention-based mod-els, it produces alignments independently at the encoder\nside using the correlation between adjacent frames. Thus, predictions can be\nmade as soon as sufficient acoustic information is received, which makes the\nmodel applicable in online cases. Besides, a small change is made to the\ndecoding stage of the encoder-decoder framework, which allows the prediction to\nexploit bidirectional contexts. We verify the ACS algorithm on a Mandarin\nspeech corpus AIShell-1, and it achieves a 31.2% CER in the online occasion,\ncompared to the 32.4% CER of the attention-based model. To fully demonstrate\nthe advantage of ACS algorithm, offline experiments are conducted, in which our\nACS model achieves an 18.7% CER, outperforming the attention-based counterpart\nwith the CER of 22.0%.","url_abs":"http://arxiv.org/abs/1808.10088v2","url_pdf":"http://arxiv.org/pdf/1808.10088v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-recognition-on-aishell-1","task":"Speech Recognition","dataset":"AISHELL-1","model":"Att","rank_in_archive_order":18,"of":18,"metrics":{"Word Error Rate (WER)":"18.7"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}