Papers › PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers

PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers

13 Feb 2024arXiv:2402.08327archive 2025-07-28

Weizhe Lin, Jingbiao Mei, Jinghong Chen, Bill Byrne

Large Multimodal Models (LMMs) excel in natural language and visual understanding but are challenged by exacting tasks such as Knowledge-based Visual Question Answering (KB-VQA) which involve the retrieval of relevant information from document collections to use in shaping answers to questions. We present an extensive training and evaluation framework, M2KR, for KB-VQA. M2KR contains a collection of vision and language tasks which we have incorporated into a single suite of benchmark tasks for training and evaluating general-purpose multi-modal retrievers. We use M2KR to develop PreFLMR, a pre-trained version of the recently developed Fine-grained Late-interaction Multi-modal Retriever (FLMR) approach to KB-VQA, and we report new state-of-the-art results across a range of tasks. We also present investigations into the scaling behaviors of PreFLMR intended to be useful in future developments in general-purpose multi-modal retrievers.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

linweizhedragon/retrieval-augmented-visual-question-answering officialmentioned on GitHubpytorchGPL-3.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Datasets

Introduced by this paper, per the archive.

M2KR

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Retrieval InfoSeek PreFLMR Recall@5 62.1 #1 of 1 Archive leaderboard report
Visual Question Answering (VQA) InfoSeek RA-VQAv2 w/ PreFLMR Accuracy 30.65 #1 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections