{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/single-shot-scene-text-retrieval","title":"Single Shot Scene Text Retrieval","arxiv_id":"1808.09044","date":"2018-08-27","proceeding":"ECCV 2018 9","authors":["Lluís Gómez","Andrés Mafla","Marçal Rusiñol","Dimosthenis Karatzas"],"abstract":"Textual information found in scene images provides high level semantic\ninformation about the image and its context and it can be leveraged for better\nscene understanding. In this paper we address the problem of scene text\nretrieval: given a text query, the system must return all images containing the\nqueried text. The novelty of the proposed model consists in the usage of a\nsingle shot CNN architecture that predicts at the same time bounding boxes and\na compact text representation of the words in them. In this way, the text based\nimage retrieval task can be casted as a simple nearest neighbor search of the\nquery text representation over the outputs of the CNN over the entire image\ndatabase. Our experiments demonstrate that the proposed architecture\noutperforms previous state-of-the-art while it offers a significant increase in\nprocessing speed.","url_abs":"http://arxiv.org/abs/1808.09044v1","url_pdf":"http://arxiv.org/pdf/1808.09044v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"single-shot-scene-text-retrieval","repo_url":"https://github.com/lluisgomez/single-shot-str","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"single-shot-scene-text-retrieval","repo_url":"https://github.com/AndresPMD/Pytorch-yolo-phoc","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"single-shot-scene-text-retrieval","repo_url":"https://github.com/DreadPiratePsyopus/Pytorch-yolo-phoc","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.09044","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}