{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/task-loss-estimation-for-sequence-prediction","title":"Task Loss Estimation for Sequence Prediction","arxiv_id":"1511.06456","date":"2015-11-19","proceeding":null,"authors":["Dzmitry Bahdanau","Dmitriy Serdyuk","Philémon Brakel","Nan Rosemary Ke","Jan Chorowski","Aaron Courville","Yoshua Bengio"],"abstract":"Often, the performance on a supervised machine learning task is evaluated\nwith a emph{task loss} function that cannot be optimized directly. Examples of\nsuch loss functions include the classification error, the edit distance and the\nBLEU score. A common workaround for this problem is to instead optimize a\nemph{surrogate loss} function, such as for instance cross-entropy or hinge\nloss. In order for this remedy to be effective, it is important to ensure that\nminimization of the surrogate loss results in minimization of the task loss, a\ncondition that we call emph{consistency with the task loss}. In this work, we\npropose another method for deriving differentiable surrogate losses that\nprovably meet this requirement. We focus on the broad class of models that\ndefine a score for every input-output pair. Our idea is that this score can be\ninterpreted as an estimate of the task loss, and that the estimation error may\nbe used as a consistent surrogate loss. A distinct feature of such an approach\nis that it defines the desirable value of the score for every input-output\npair. We use this property to design specialized surrogate losses for\nEncoder-Decoder models often used for sequence prediction tasks. In our\nexperiment, we benchmark on the task of speech recognition. Using a new\nsurrogate loss instead of cross-entropy to train an Encoder-Decoder speech\nrecognizer brings a significant ~13% relative improvement in terms of Character\nError Rate (CER) in the case when no extra corpora are used for language\nmodeling.","url_abs":"http://arxiv.org/abs/1511.06456v4","url_pdf":"http://arxiv.org/pdf/1511.06456v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"task-loss-estimation-for-sequence-prediction","repo_url":"https://github.com/rizar/attention-lvcsr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.06456","atlas_url":"https://app.syntology.ai/?focus=1511.06456","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1511.06456"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rizar/attention-lvcsr","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f20f6bc83bfeae21","entry":"add_raw_text","repo":"rizar/attention-lvcsr","repo_kind":"official","path":"bin/kaldi2fuel.py","file_url":"https://github.com/rizar/attention-lvcsr/blob/HEAD/bin/kaldi2fuel.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f20f6bc83bfeae21"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}