{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-text-correction","title":"Visual Text Correction","arxiv_id":"1801.01967","date":"2018-01-06","proceeding":"ECCV 2018 9","authors":["Amir Mazaheri","Mubarak Shah"],"abstract":"Videos, images, and sentences are mediums that can express the same\nsemantics. One can imagine a picture by reading a sentence or can describe a\nscene with some words. However, even small changes in a sentence can cause a\nsignificant semantic inconsistency with the corresponding video/image. For\nexample, by changing the verb of a sentence, the meaning may drastically\nchange. There have been many efforts to encode a video/sentence and decode it\nas a sentence/video. In this research, we study a new scenario in which both\nthe sentence and the video are given, but the sentence is inaccurate. A\nsemantic inconsistency between the sentence and the video or between the words\nof a sentence can result in an inaccurate description. This paper introduces a\nnew problem, called Visual Text Correction (VTC), i.e., finding and replacing\nan inaccurate word in the textual description of a video. We propose a deep\nnetwork that can simultaneously detect an inaccuracy in a sentence, and fix it\nby replacing the inaccurate word(s). Our method leverages the semantic\ninterdependence of videos and words, as well as the short-term and long-term\nrelations of the words in a sentence. In our formulation, part of a visual\nfeature vector for every single word is dynamically selected through a gating\nprocess. Furthermore, to train and evaluate our model, we propose an approach\nto automatically construct a large dataset for VTC problem. Our experiments and\nperformance analysis demonstrates that the proposed method provides very good\nresults and also highlights the general challenges in solving the VTC problem.\nTo the best of our knowledge, this work is the first of its kind for the Visual\nText Correction task.","url_abs":"http://arxiv.org/abs/1801.01967v3","url_pdf":"http://arxiv.org/pdf/1801.01967v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-text-correction","repo_url":"https://github.com/amirmazaheri1990/Visual-Text-Correction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"grammatical-error-correction","task_name":"Grammatical Error Correction"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"visual-text-correction","task_name":"Visual Text Correction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1801.01967","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}