{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lamra-large-multimodal-model-as-your-advanced","title":"LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant","arxiv_id":"2412.01720","date":"2024-12-02","proceeding":"CVPR 2025 1","authors":["Yikun Liu","Pingan Chen","Jiayin Cai","XiaoLong Jiang","Yao Hu","Jiangchao Yao","Yanfeng Wang","Weidi Xie"],"abstract":"With the rapid advancement of multimodal information retrieval, increasingly complex retrieval tasks have emerged. Existing methods predominately rely on task-specific fine-tuning of vision-language models, often those trained with image-text contrastive learning. In this paper, we explore the possibility of re-purposing generative Large Multimodal Models (LMMs) for retrieval. This approach enables unifying all retrieval tasks under the same formulation and, more importantly, allows for extrapolation towards unseen retrieval tasks without additional training. Our contributions can be summarised in the following aspects: (i) We introduce LamRA, a versatile framework designed to empower LMMs with sophisticated retrieval and reranking capabilities. (ii) For retrieval, we adopt a two-stage training strategy comprising language-only pre-training and multimodal instruction tuning to progressively enhance LMM's retrieval performance. (iii) For reranking, we employ joint training for both pointwise and listwise reranking, offering two distinct ways to further boost the retrieval performance. (iv) Extensive experimental results underscore the efficacy of our method in handling more than ten retrieval tasks, demonstrating robust performance in both supervised and zero-shot settings, including scenarios involving previously unseen retrieval tasks.","url_abs":"https://arxiv.org/abs/2412.01720v1","url_pdf":"https://arxiv.org/pdf/2412.01720v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lamra-large-multimodal-model-as-your-advanced","repo_url":"https://github.com/Code-kunkun/LamRA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"reranking","task_name":"Reranking"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[{"method_slug":"adopt","method_name":"ADOPT"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2412.01720","atlas_url":"https://app.syntology.ai/?focus=2412.01720","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2412.01720"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Code-kunkun/LamRA","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_honours":3,"ran":2},"by_repo_kind":{"listed":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"6e45201fa27cb24a","entry":"ceil_by_factor","repo":"Code-kunkun/LamRA","repo_kind":"listed","path":"collators/qwen2_vision_process.py","file_url":"https://github.com/Code-kunkun/LamRA/blob/HEAD/collators/qwen2_vision_process.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6e45201fa27cb24a"}},{"code_sha256_prefix":"60134dd1f6eb9025","entry":"extract_inputs","repo":"Code-kunkun/LamRA","repo_kind":"listed","path":"collators/mbeir_rerank.py","file_url":"https://github.com/Code-kunkun/LamRA/blob/HEAD/collators/mbeir_rerank.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"60134dd1f6eb9025"}},{"code_sha256_prefix":"db7c18eeb2aa3486","entry":"find_all_linear_names","repo":"Code-kunkun/LamRA","repo_kind":"listed","path":"utils.py","file_url":"https://github.com/Code-kunkun/LamRA/blob/HEAD/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"db7c18eeb2aa3486"}},{"code_sha256_prefix":"8155263d7ff19bb3","entry":"floor_by_factor","repo":"Code-kunkun/LamRA","repo_kind":"listed","path":"collators/qwen2_vision_process.py","file_url":"https://github.com/Code-kunkun/LamRA/blob/HEAD/collators/qwen2_vision_process.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8155263d7ff19bb3"}},{"code_sha256_prefix":"e252767324188623","entry":"round_by_factor","repo":"Code-kunkun/LamRA","repo_kind":"listed","path":"collators/qwen2_vision_process.py","file_url":"https://github.com/Code-kunkun/LamRA/blob/HEAD/collators/qwen2_vision_process.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e252767324188623"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}