{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/object-captioning-and-retrieval-with-natural","title":"Object Captioning and Retrieval with Natural Language","arxiv_id":"1803.06152","date":"2018-03-16","proceeding":null,"authors":["Anh Nguyen","Thanh-Toan Do","Ian Reid","Darwin G. Caldwell","Nikos G. Tsagarakis"],"abstract":"We address the problem of jointly learning vision and language to understand\nthe object in a fine-grained manner. The key idea of our approach is the use of\nobject descriptions to provide the detailed understanding of an object. Based\non this idea, we propose two new architectures to solve two related problems:\nobject captioning and natural language-based object retrieval. The goal of the\nobject captioning task is to simultaneously detect the object and generate its\nassociated description, while in the object retrieval task, the goal is to\nlocalize an object given an input query. We demonstrate that both problems can\nbe solved effectively using hybrid end-to-end CNN-LSTM networks. The\nexperimental results on our new challenging dataset show that our methods\noutperform recent methods by a fair margin, while providing a detailed\nunderstanding of the object and having fast inference time. The source code\nwill be made available.","url_abs":"http://arxiv.org/abs/1803.06152v1","url_pdf":"http://arxiv.org/pdf/1803.06152v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"object-captioning-and-retrieval-with-natural","repo_url":"https://github.com/nqanh/object_captioning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1803.06152","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}