{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-re-ranking-with-natural-language","title":"Visual Re-ranking with Natural Language Understanding for Text Spotting","arxiv_id":"1810.12738","date":"2018-10-29","proceeding":null,"authors":["Ahmed Sabir","Francesc Moreno-Noguer","Lluís Padró"],"abstract":"Many scene text recognition approaches are based on purely visual information\nand ignore the semantic relation between scene and text. In this paper, we\ntackle this problem from natural language processing perspective to fill the\ngap between language and vision. We propose a post-processing approach to\nimprove scene text recognition accuracy by using occurrence probabilities of\nwords (unigram language model), and the semantic correlation between scene and\ntext. For this, we initially rely on an off-the-shelf deep neural network,\nalready trained with a large amount of data, which provides a series of text\nhypotheses per input image. These hypotheses are then re-ranked using word\nfrequencies and semantic relatedness with objects or scenes in the image. As a\nresult of this combination, the performance of the original network is boosted\nwith almost no additional cost. We validate our approach on ICDAR'17 dataset.","url_abs":"http://arxiv.org/abs/1810.12738v1","url_pdf":"http://arxiv.org/pdf/1810.12738v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-re-ranking-with-natural-language","repo_url":"https://github.com/ahmedssabir/Visual-Semantic-Relatedness-with-Word-Embedding","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"visual-re-ranking-with-natural-language","repo_url":"https://github.com/ahmedssabir/dataset","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"visual-re-ranking-with-natural-language","repo_url":"https://github.com/ahmedssabir/Textual-Visual-Semantic-Dataset-for-Text-Spotting","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"},{"task_slug":"re-ranking","task_name":"Re-Ranking"},{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"},{"task_slug":"text-spotting","task_name":"Text Spotting"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}