{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dialog-based-interactive-image-retrieval","title":"Dialog-based Interactive Image Retrieval","arxiv_id":"1805.00145","date":"2018-05-01","proceeding":"NeurIPS 2018 12","authors":["Xiaoxiao Guo","Hui Wu","Yu Cheng","Steven Rennie","Gerald Tesauro","Rogerio Schmidt Feris"],"abstract":"Existing methods for interactive image retrieval have demonstrated the merit\nof integrating user feedback, improving retrieval results. However, most\ncurrent systems rely on restricted forms of user feedback, such as binary\nrelevance responses, or feedback based on a fixed set of relative attributes,\nwhich limits their impact. In this paper, we introduce a new approach to\ninteractive image search that enables users to provide feedback via natural\nlanguage, allowing for more natural and effective interaction. We formulate the\ntask of dialog-based interactive image retrieval as a reinforcement learning\nproblem, and reward the dialog system for improving the rank of the target\nimage during each dialog turn. To mitigate the cumbersome and costly process of\ncollecting human-machine conversations as the dialog system learns, we train\nour system with a user simulator, which is itself trained to describe the\ndifferences between target and candidate images. The efficacy of our approach\nis demonstrated in a footwear retrieval application. Experiments on both\nsimulated and real-world data show that 1) our proposed learning framework\nachieves better accuracy than other supervised and reinforcement learning\nbaselines and 2) user feedback based on natural language rather than\npre-specified attributes leads to more effective retrieval results, and a more\nnatural and expressive communication interface.","url_abs":"http://arxiv.org/abs/1805.00145v3","url_pdf":"http://arxiv.org/pdf/1805.00145v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dialog-based-interactive-image-retrieval","repo_url":"https://github.com/XiaoxiaoGuo/fashion-retrieval","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"visual-dialogue","task_name":"Visual Dialog"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1805.00145","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}