{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/show-translate-and-tell","title":"Show, Translate and Tell","arxiv_id":"1903.06275","date":"2019-03-14","proceeding":null,"authors":["Dheeraj Peri","Shagan Sah","Raymond Ptucha"],"abstract":"Humans have an incredible ability to process and understand information from\nmultiple sources such as images, video, text, and speech. Recent success of\ndeep neural networks has enabled us to develop algorithms which give machines\nthe ability to understand and interpret this information. There is a need to\nboth broaden their applicability and develop methods which correlate visual\ninformation along with semantic content. We propose a unified model which\njointly trains on images and captions, and learns to generate new captions\ngiven either an image or a caption query. We evaluate our model on three\ndifferent tasks namely cross-modal retrieval, image captioning, and sentence\nparaphrasing. Our model gains insight into cross-modal vector embeddings,\ngeneralizes well on multiple tasks and is competitive to state of the art\nmethods on retrieval.","url_abs":"http://arxiv.org/abs/1903.06275v1","url_pdf":"http://arxiv.org/pdf/1903.06275v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"show-translate-and-tell","repo_url":"https://github.com/peri044/STT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}