Papers › Attention-Based Models for Speech Recognition

Attention-Based Models for Speech Recognition

24 Jun 2015NeurIPS 2015 12arXiv:1506.07503archive 2025-07-28

Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, Yoshua Bengio

Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend the attention-mechanism with features needed for speech recognition. We show that while an adaptation of the model used for machine translation in reaches a competitive 18.7% phoneme error rate (PER) on the TIMIT phoneme recognition task, it can only be applied to utterances which are roughly as long as the ones it was trained on. We offer a qualitative explanation of this failure and propose a novel and generic method of adding location-awareness to the attention mechanism to alleviate this issue. The new method yields a model that is robust to long inputs and achieves 18% PER in single utterances and 20% in 10-times longer (repeated) utterances. Finally, we propose a change to the at- tention mechanism that prevents it from concentrating too much on single frames, which further reduces PER to 17.6% level.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

14 repositories listed; official and paper-mentioned ones first.

Alexander-H-Liu/End-to-end-ASR-Pytorch mentioned on GitHubpytorchMIT report
CKRC24/Listen-and-Translate mentioned on GitHubtf report
biyoml/End-to-End-Mandarin-ASR mentioned on GitHubpytorch report
biyoml/Pytorch-End-to-End-ASR-on-TIMIT mentioned on GitHubpytorch report
jackjhliu/End-to-End-Mandarin-ASR mentioned on GitHubpytorch report
mnm-rnd/elsa-voice-asr mentioned on GitHubpytorchMIT report
msalhab96/SpeeQ mentioned on GitHubpytorch report
neil-zeng/asr mentioned on GitHubpytorchMIT report
s3prl/End-to-end-ASR-Pytorch mentioned on GitHubpytorchMIT report
sooftware/End-to-end-Speech-Recognition mentioned on GitHubpytorch report
sooftware/OpenSpeech mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Machine TranslationPhoneme RecognitionSpeech RecognitionTranslation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Recognition TIMIT Bi-RNN + Attention Percentage error 17.6 #17 of 22 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Introduced by this paper: Location Sensitive Attention

Location Sensitive AttentionTanh Activation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections