{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/engaging-image-chat-modeling-personality-in","title":"Image Chat: Engaging Grounded Conversations","arxiv_id":"1811.00945","date":"2018-11-02","proceeding":null,"authors":["Kurt Shuster","Samuel Humeau","Antoine Bordes","Jason Weston"],"abstract":"To achieve the long-term goal of machines being able to engage humans in conversation, our models should captivate the interest of their speaking partners. Communication grounded in images, whereby a dialogue is conducted based on a given photo, is a setup naturally appealing to humans (Hu et al., 2014). In this work we study large-scale architectures and datasets for this goal. We test a set of neural architectures using state-of-the-art image and text representations, considering various ways to fuse the components. To test such models, we collect a dataset of grounded human-human conversations, where speakers are asked to play roles given a provided emotional mood or style, as the use of such traits is also a key factor in engagingness (Guo et al., 2019). Our dataset, Image-Chat, consists of 202k dialogues over 202k images using 215 possible style traits. Automatic metrics and human evaluations of engagingness show the efficacy of our approach; in particular, we obtain state-of-the-art performance on the existing IGC task, and our best performing model is almost on par with humans on the Image-Chat test set (preferred 47.7% of the time).","url_abs":"https://arxiv.org/abs/1811.00945v2","url_pdf":"https://arxiv.org/pdf/1811.00945v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"engaging-image-chat-modeling-personality-in","repo_url":"https://github.com/Alenush/sirius-spring2021-image2chat","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"engaging-image-chat-modeling-personality-in","repo_url":"https://github.com/facebookresearch/ParlAI","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"engaging-image-chat-modeling-personality-in","repo_url":"https://github.com/joe-prog/https-github.com-facebookresearch-ParlAI","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"gone","observed_at":"2026-09-18","how":"tree_404+repo_404"}}],"tasks":[{"task_slug":"text-retrieval","task_name":"Text Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-retrieval-on-image-chat","task":"Text Retrieval","dataset":"Image-Chat","model":"TransResNet","rank_in_archive_order":2,"of":3,"metrics":{"R@1":"50.3","R@5":"75.4","Sum(R@1,5)":"125.7"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.00945","atlas_url":"https://app.syntology.ai/?focus=1811.00945","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}