{"url":"/dataset/m2kr","name":"M2KR","full_name":"Multi-task Multi-modal Knowledge Retrieval","description_markdown":"The **M2KR** is a collection of datasets designed for training and evaluating general-purpose vision-language retrievers. These datasets are released in **Huggingface Dataset format** and cover various retrieval tasks. Let's delve into the details:\r\n\r\n1. **Image to Text (I2T) retrieval**: This task involves retrieving relevant textual descriptions given an input image.\r\n2. **Question to Text (Q2T) retrieval**: Here, the goal is to retrieve relevant text passages based on a given question.\r\n3. **Image & Question to Text (IQ2T) retrieval**: This task combines both image and question inputs to retrieve relevant textual information.\r\n\r\nThe M2KR benchmark comprises nine datasets, each tailored for specific tasks. Some of these datasets include:\r\n\r\n- **WIT (Web Image Text)**: A dataset for I2T retrieval.\r\n- **IGLUE (Image-Grounded Language Understanding Evaluation)**: Used for Q2T retrieval.\r\n- **KVQA (Knowledge Visual Question Answering)**: Relevant for IQ2T retrieval.\r\n- **CC3M (Common Crawl 3 Million)**: Another dataset for IQ2T retrieval.\r\n- **OVEN (Open Vision and Language Evaluation)**: Used in IQ2T retrieval.\r\n- **LLaVA (Large-scale Language-Visual Association)**: Relevant for I2T retrieval.\r\n- **OKVQA (Open Knowledge Visual Question Answering)**: Used in IQ2T retrieval.\r\n- **Infoseek**: A dataset for I2T retrieval.\r\n- **E-VQA (English Visual Question Answering)**: Relevant for IQ2T retrieval.\r\n\r\nThese datasets enable researchers to develop and evaluate vision-language models, and they play a crucial role in advancing the field of multimodal understanding and retrieval¹².\r\n\r\n(1) M2KR Benchmark Datasets - GitHub. https://github.com/LinWeizheDragon/FLMR/blob/main/docs/Datasets.md.\r\n(2) arXiv:2402.08327v1 [cs.CL] 13 Feb 2024. https://arxiv.org/pdf/2402.08327.pdf.\r\n(3) Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers. https://preflmr.github.io/.\r\n(4) undefined. https://avatars.githubusercontent.com/u/33350454?v=4.\r\n(5) undefined. https://github.com/LinWeizheDragon/FLMR/blob/main/docs/Datasets.md?raw=true.\r\n(6) undefined. https://desktop.github.com.\r\n(7) undefined. https://github.com/LinWeizheDragon/FLMR/raw/main/docs/Datasets.md.","description_withheld":null,"homepage":"https://github.com/LinWeizheDragon/FLMR","introduced_date":"2024-02-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/preflmr-scaling-up-fine-grained-late","title":"PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers","first_author":"Weizhe Lin","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["M2KR"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}