{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/language-models-as-knowledge-bases-for-visual","title":"Language Models as Knowledge Bases for Visual Word Sense Disambiguation","arxiv_id":"2310.01960","date":"2023-10-03","proceeding":null,"authors":["Anastasia Kritharoula","Maria Lymperaiou","Giorgos Stamou"],"abstract":"Visual Word Sense Disambiguation (VWSD) is a novel challenging task that lies between linguistic sense disambiguation and fine-grained multimodal retrieval. The recent advancements in the development of visiolinguistic (VL) transformers suggest some off-the-self implementations with encouraging results, which however we argue that can be further improved. To this end, we propose some knowledge-enhancement techniques towards improving the retrieval performance of VL transformers via the usage of Large Language Models (LLMs) as Knowledge Bases. More specifically, knowledge stored in LLMs is retrieved with the help of appropriate prompts in a zero-shot manner, achieving performance advancements. Moreover, we convert VWSD to a purely textual question-answering (QA) problem by considering generated image captions as multiple-choice candidate answers. Zero-shot and few-shot prompting strategies are leveraged to explore the potential of such a transformation, while Chain-of-Thought (CoT) prompting in the zero-shot setting is able to reveal the internal reasoning steps an LLM follows to select the appropriate candidate. In total, our presented approach is the first one to analyze the merits of exploiting knowledge stored in LLMs in different ways to solve WVSD.","url_abs":"https://arxiv.org/abs/2310.01960v1","url_pdf":"https://arxiv.org/pdf/2310.01960v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"language-models-as-knowledge-bases-for-visual","repo_url":"https://github.com/anastasiakrith/llm-for-vwsd","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"word-sense-disambiguation","task_name":"Word Sense Disambiguation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2310.01960","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.01960"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/anastasiakrith/llm-for-vwsd","reach":{"status":"ok"}}],"summary":{"ran":3,"ran_honours":1,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"e4a5fab751465606","entry":"CoT_prompt","repo":"anastasiakrith/llm-for-vwsd","repo_kind":"official","path":"modules/prompts.py","file_url":"https://github.com/anastasiakrith/llm-for-vwsd/blob/HEAD/modules/prompts.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e4a5fab751465606"}},{"code_sha256_prefix":"f3ce5fa0f9d664c2","entry":"get_k_default_value","repo":"anastasiakrith/llm-for-vwsd","repo_kind":"official","path":"qa_retrieval_eval.py","file_url":"https://github.com/anastasiakrith/llm-for-vwsd/blob/HEAD/qa_retrieval_eval.py","link_basis":"plan_row","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f3ce5fa0f9d664c2"}},{"code_sha256_prefix":"5ee88d1ab88ef45c","entry":"no_CoT_prompt","repo":"anastasiakrith/llm-for-vwsd","repo_kind":"official","path":"modules/prompts.py","file_url":"https://github.com/anastasiakrith/llm-for-vwsd/blob/HEAD/modules/prompts.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5ee88d1ab88ef45c"}},{"code_sha256_prefix":"d8e8c3be701c6eb7","entry":"think_prompt","repo":"anastasiakrith/llm-for-vwsd","repo_kind":"official","path":"modules/prompts.py","file_url":"https://github.com/anastasiakrith/llm-for-vwsd/blob/HEAD/modules/prompts.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d8e8c3be701c6eb7"}},{"code_sha256_prefix":"a2fb77ded7a905fa","entry":"clean_captions","repo":"anastasiakrith/llm-for-vwsd","repo_kind":"official","path":"modules/captioners.py","file_url":"https://github.com/anastasiakrith/llm-for-vwsd/blob/HEAD/modules/captioners.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a2fb77ded7a905fa"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}