{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-attention-networks-for-connectionist","title":"Self-Attention Networks for Connectionist Temporal Classification in Speech Recognition","arxiv_id":"1901.10055","date":"2019-01-22","proceeding":null,"authors":["Julian Salazar","Katrin Kirchhoff","Zhiheng Huang"],"abstract":"The success of self-attention in NLP has led to recent applications in\nend-to-end encoder-decoder architectures for speech recognition. Separately,\nconnectionist temporal classification (CTC) has matured as an alignment-free,\nnon-autoregressive approach to sequence transduction, either by itself or in\nvarious multitask and decoding frameworks. We propose SAN-CTC, a deep, fully\nself-attentional network for CTC, and show it is tractable and competitive for\nend-to-end speech recognition. SAN-CTC trains quickly and outperforms existing\nCTC models and most encoder-decoder models, with character error rates (CERs)\nof 4.7% in 1 day on WSJ eval92 and 2.8% in 1 week on LibriSpeech test-clean,\nwith a fixed architecture and one GPU. Similar improvements hold for WERs after\nLM decoding. We motivate the architecture for speech, evaluate position and\ndownsampling approaches, and explore how label alphabets (character, phoneme,\nsubword) affect attention heads and performance.","url_abs":"http://arxiv.org/abs/1901.10055v2","url_pdf":"http://arxiv.org/pdf/1901.10055v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"self-attention-networks-for-connectionist","repo_url":"https://github.com/aaaceo890/Attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":null,"task_name":"Position"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.10055","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}