{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/preflmr-scaling-up-fine-grained-late","title":"PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers","arxiv_id":"2402.08327","date":"2024-02-13","proceeding":null,"authors":["Weizhe Lin","Jingbiao Mei","Jinghong Chen","Bill Byrne"],"abstract":"Large Multimodal Models (LMMs) excel in natural language and visual understanding but are challenged by exacting tasks such as Knowledge-based Visual Question Answering (KB-VQA) which involve the retrieval of relevant information from document collections to use in shaping answers to questions. We present an extensive training and evaluation framework, M2KR, for KB-VQA. M2KR contains a collection of vision and language tasks which we have incorporated into a single suite of benchmark tasks for training and evaluating general-purpose multi-modal retrievers. We use M2KR to develop PreFLMR, a pre-trained version of the recently developed Fine-grained Late-interaction Multi-modal Retriever (FLMR) approach to KB-VQA, and we report new state-of-the-art results across a range of tasks. We also present investigations into the scaling behaviors of PreFLMR intended to be useful in future developments in general-purpose multi-modal retrievers.","url_abs":"https://arxiv.org/abs/2402.08327v2","url_pdf":"https://arxiv.org/pdf/2402.08327v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"preflmr-scaling-up-fine-grained-late","repo_url":"https://github.com/linweizhedragon/retrieval-augmented-visual-question-answering","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[{"slug":"m2kr","name":"M2KR","full_name":"Multi-task Multi-modal Knowledge Retrieval"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/retrieval-on-infoseek","task":"Retrieval","dataset":"InfoSeek","model":"PreFLMR","rank_in_archive_order":1,"of":1,"metrics":{"Recall@5":"62.1"},"uses_additional_data":true},{"leaderboard":"/sota/visual-question-answering-vqa-on-infoseek","task":"Visual Question Answering (VQA)","dataset":"InfoSeek","model":"RA-VQAv2 w/ PreFLMR","rank_in_archive_order":1,"of":7,"metrics":{"Accuracy":"30.65"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.08327","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}