{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/resee-responding-through-seeing-fine-grained","title":"ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain Dialogue","arxiv_id":"2305.13602","date":"2023-05-23","proceeding":null,"authors":["Haoqin Tu","Yitong Li","Fei Mi","Zhongliang Yang"],"abstract":"Incorporating visual knowledge into text-only dialogue systems has become a potential direction to imitate the way humans think, imagine, and communicate. However, existing multimodal dialogue systems are either confined by the scale and quality of available datasets or the coarse concept of visual knowledge. To address these issues, we provide a new paradigm of constructing multimodal dialogues as well as two datasets extended from text-only dialogues under such paradigm (ReSee-WoW, ReSee-DD). We propose to explicitly split the visual knowledge into finer granularity (``turn-level'' and ``entity-level''). To further boost the accuracy and diversity of augmented visual information, we retrieve them from the Internet or a large image dataset. To demonstrate the superiority and universality of the provided visual knowledge, we propose a simple but effective framework ReSee to add visual representation into vanilla dialogue models by modality concatenations. We also conduct extensive experiments and ablations w.r.t. different model configurations and visual knowledge settings. Empirical, encouraging results not only demonstrate the effectiveness of introducing visual knowledge at both entity and turn level but also verify the proposed model ReSee outperforms several state-of-the-art methods on automatic and human evaluations. By leveraging text and vision knowledge, ReSee can produce informative responses with real-world visual concepts. Our code is available at https://github.com/ImKeTT/ReSee.","url_abs":"https://arxiv.org/abs/2305.13602v2","url_pdf":"https://arxiv.org/pdf/2305.13602v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"resee-responding-through-seeing-fine-grained","repo_url":"https://github.com/imkett/resee","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2305.13602","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2305.13602"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/ImKeTT/ReSee","reach":null}],"summary":{"ran":1,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"eecafe8e4da9e0ac","entry":"GPT2ModelForMultiModal","repo":"ImKeTT/ReSee","repo_kind":"official","path":"model/base.py","file_url":"https://github.com/ImKeTT/ReSee/blob/HEAD/model/base.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"eecafe8e4da9e0ac"}},{"code_sha256_prefix":"098f97722079692d","entry":"GPT2LMforMMDialog","repo":"ImKeTT/ReSee","repo_kind":"official","path":"model/base.py","file_url":"https://github.com/ImKeTT/ReSee/blob/HEAD/model/base.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"098f97722079692d"}},{"code_sha256_prefix":"3fd446cb7344b0c0","entry":"MMGPT2Block","repo":"ImKeTT/ReSee","repo_kind":"official","path":"model/base.py","file_url":"https://github.com/ImKeTT/ReSee/blob/HEAD/model/base.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3fd446cb7344b0c0"}},{"code_sha256_prefix":"194896a8bb5370d2","entry":"UnmaskedAttention","repo":"ImKeTT/ReSee","repo_kind":"official","path":"model/base.py","file_url":"https://github.com/ImKeTT/ReSee/blob/HEAD/model/base.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"194896a8bb5370d2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}