{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/investigating-the-effects-of-word","title":"Investigating the Effects of Word Substitution Errors on Sentence Embeddings","arxiv_id":"1811.07021","date":"2018-11-16","proceeding":null,"authors":["Rohit Voleti","Julie M. Liss","Visar Berisha"],"abstract":"A key initial step in several natural language processing (NLP) tasks\ninvolves embedding phrases of text to vectors of real numbers that preserve\nsemantic meaning. To that end, several methods have been recently proposed with\nimpressive results on semantic similarity tasks. However, all of these\napproaches assume that perfect transcripts are available when generating the\nembeddings. While this is a reasonable assumption for analysis of written text,\nit is limiting for analysis of transcribed text. In this paper we investigate\nthe effects of word substitution errors, such as those coming from automatic\nspeech recognition errors (ASR), on several state-of-the-art sentence embedding\nmethods. To do this, we propose a new simulator that allows the experimenter to\ninduce ASR-plausible word substitution errors in a corpus at a desired word\nerror rate. We use this simulator to evaluate the robustness of several\nsentence embedding methods. Our results show that pre-trained neural sentence\nencoders are both robust to ASR errors and perform well on textual similarity\ntasks after errors are introduced. Meanwhile, unweighted averages of word\nvectors perform well with perfect transcriptions, but their performance\ndegrades rapidly on textual similarity tasks for text with word substitution\nerrors.","url_abs":"http://arxiv.org/abs/1811.07021v2","url_pdf":"http://arxiv.org/pdf/1811.07021v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"investigating-the-effects-of-word","repo_url":"https://github.com/rvoleti89/icaasp-2019-word-subs","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentence-embedding","task_name":"Sentence Embedding"},{"task_slug":"sentence-embeddings","task_name":"Sentence Embeddings"},{"task_slug":"sentence-embedding-1","task_name":"Sentence-Embedding"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}