{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scene-text-oriented-reffering-expression","title":"Scene-Text Oriented Reffering Expression Comprehension","arxiv_id":null,"date":"2022-11-04","proceeding":"2023 2022 11","authors":["Yuqi Bu","Liuwu Li","Jiayuan Xie","Qiong Liu","Yi Cai","Qingbao Huang","Qing Li"],"abstract":"Abstract—Referring expression comprehension (REC) aims to\r\nidentify and locate a specific object in visual scenes referred\r\nto by a natural language expression. Existing studies of REC\r\nonly focus on basic visual attributes and neglect scene text.\r\nSince scene text has the functions of object identification and\r\ndisambiguation, it is naturally and frequently used to refer to\r\nobjects. However, existing methods do not explicitly recognize text\r\nin images and fail to align scene text mentioned in expressions\r\nwith the text shown in images, resulting in object localization\r\nerrors. This study takes the first step toward addressing these\r\nlimitations. First, we introduce a new task called scene-text\r\noriented referring expression comprehension, which aims to align\r\nvisual cues and textual semantics of scene text with referring\r\nexpressions and visual contents. Second, we propose a scene text\r\nawareness network that can bridge the gap between texts from\r\ntwo modalities by grounding visual representations of expressioncorrelated scene texts. Specifically, we propose a correlated text\r\nextraction module to solve the problem of lacking semantic understanding, and a correlated region activation module to address\r\nthe fixed alignment problem and absent alignment problem.\r\nThese modules ensure that the proposed method focuses on local\r\nregions that are most relevant to scene text, thus mitigating the\r\nmisalignment of scene text with irrelevant regions. Third, to\r\nconduct quantitative evaluations, we establish a new benchmark\r\ndataset called RefText. Experimental results demonstrate that the\r\nproposed method can effectively comprehend scene-text oriented\r\nreferring expressions and achieves excellent performance.\r\nIndex Terms—Referring expression comprehension, scene text\r\nrepresentation, multimodal alignment.","url_abs":"https://ieeexplore.ieee.org/abstract/document/9939075","url_pdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9939075&tag=1","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scene-text-oriented-reffering-expression","repo_url":"https://github.com/2023-MindSpore-4/Code4/tree/main/stan","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"object-localization","task_name":"Object Localization"},{"task_slug":"referring-expression","task_name":"Referring Expression"},{"task_slug":"referring-expression-comprehension","task_name":"Referring Expression Comprehension"}],"methods":[{"method_slug":"align","method_name":"ALIGN"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}