{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/phonetically-oriented-word-error-alignment","title":"Phonetically-Oriented Word Error Alignment for Speech Recognition Error Analysis in Speech Translation","arxiv_id":"1904.11024","date":"2019-04-24","proceeding":null,"authors":["Nicholas Ruiz","Marcello Federico"],"abstract":"We propose a variation to the commonly used Word Error Rate (WER) metric for\nspeech recognition evaluation which incorporates the alignment of phonemes, in\nthe absence of time boundary information. After computing the Levenshtein\nalignment on words in the reference and hypothesis transcripts, spans of\nadjacent errors are converted into phonemes with word and syllable boundaries\nand a phonetic Levenshtein alignment is performed. The aligned phonemes are\nrecombined into aligned words that adjust the word alignment labels in each\nerror region. We demonstrate that our Phonetically-Oriented Word Error Rate\n(POWER) yields similar scores to WER with the added advantages of better word\nalignments and the ability to capture one-to-many word alignments corresponding\nto homophonic errors in speech recognition hypotheses. These improved\nalignments allow us to better trace the impact of Levenshtein error types on\ndownstream tasks such as speech translation.","url_abs":"http://arxiv.org/abs/1904.11024v1","url_pdf":"http://arxiv.org/pdf/1904.11024v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"phonetically-oriented-word-error-alignment","repo_url":"https://github.com/NickRuiz/power-asr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"word-alignment","task_name":"Word Alignment"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}