{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2512-05119","title":"RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering","arxiv_id":"2512.05119","date":"2025-10-11","proceeding":"NeurIPS","authors":["Rongyang Zhang","Yuqing Huang","Chengqiang Lu","Qimeng Wang","Yan Gao","Yi Wu","Yao Hu","Yin Xu","Wei Wang","Hao Wang","Enhong Chen"],"abstract":"In real-world scenarios, providing user queries with visually enhanced responses can considerably benefit understanding and memory, underscoring the great value of interleaved image-text generation. Despite recent progress, like the visual autoregressive model that unifies text and image processing in a single transformer architecture, generating high-quality interleaved content remains challenging. Moreover, evaluations of these interleaved sequences largely remain underexplored, with existing benchmarks often limited by unimodal metrics that inadequately assess the intricacies of combined image-text outputs. To address these issues, we present RAG-IGBench, a thorough benchmark designed specifically to evaluate the task of Interleaved Generation based on Retrieval-Augmented Generation (RAG-IG) in open-domain question answering. RAG-IG integrates multimodal large language models (MLLMs) with retrieval mechanisms, enabling the models to access external image-text information for generating coherent multimodal content. Distinct from previous datasets, RAG-IGBench draws on the latest publicly available content from social platforms and introduces innovative evaluation metrics that measure the quality of text and images, as well as their consistency. Through extensive experiments with state-of-the-art MLLMs (both open-source and proprietary) on RAG-IGBench, we provide an in-depth analysis examining the capabilities and limitations of these models. Additionally, we validate our evaluation metrics by demonstrating their high correlation with human assessments. Models fine-tuned on RAG-IGBench's training set exhibit improved performance across multiple benchmarks, confirming both the quality and practical utility of our dataset. Our benchmark is available at https://github.com/USTC-StarTeam/RAG-IGBench.","url_abs":"https://arxiv.org/abs/2512.05119","url_pdf":"https://arxiv.org/pdf/2512.05119","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2512.05119","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2512.05119"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/USTC-StarTeam/RAG-IGBench","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran_draft_wrong":1,"unverified":7},"by_repo_kind":{"found_in_text":{"samples":8,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f41cb1a19b154297","entry":"encode_image","repo":"USTC-StarTeam/RAG-IGBench","repo_kind":"found_in_text","path":"model_generation/claude.py","file_url":"https://github.com/USTC-StarTeam/RAG-IGBench/blob/HEAD/model_generation/claude.py","link_basis":"plan_row","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f41cb1a19b154297"}},{"code_sha256_prefix":"7c3de639a96dcf10","entry":"convert_prompt","repo":"USTC-StarTeam/RAG-IGBench","repo_kind":"found_in_text","path":"model_generation/claude.py","file_url":"https://github.com/USTC-StarTeam/RAG-IGBench/blob/HEAD/model_generation/claude.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"7c3de639a96dcf10"}},{"code_sha256_prefix":"84245264a71f79ac","entry":"convert_prompt","repo":"USTC-StarTeam/RAG-IGBench","repo_kind":"found_in_text","path":"model_generation/gemini.py","file_url":"https://github.com/USTC-StarTeam/RAG-IGBench/blob/HEAD/model_generation/gemini.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"84245264a71f79ac"}},{"code_sha256_prefix":"68380f6694c4d9ae","entry":"convert_prompt","repo":"USTC-StarTeam/RAG-IGBench","repo_kind":"found_in_text","path":"model_generation/gpt4o.py","file_url":"https://github.com/USTC-StarTeam/RAG-IGBench/blob/HEAD/model_generation/gpt4o.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"68380f6694c4d9ae"}},{"code_sha256_prefix":"22707879c5a43031","entry":"convert_prompt","repo":"USTC-StarTeam/RAG-IGBench","repo_kind":"found_in_text","path":"model_generation/qwen2vl.py","file_url":"https://github.com/USTC-StarTeam/RAG-IGBench/blob/HEAD/model_generation/qwen2vl.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"22707879c5a43031"}},{"code_sha256_prefix":"6c277445442e0177","entry":"download_image","repo":"USTC-StarTeam/RAG-IGBench","repo_kind":"found_in_text","path":"download_images.py","file_url":"https://github.com/USTC-StarTeam/RAG-IGBench/blob/HEAD/download_images.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6c277445442e0177"}},{"code_sha256_prefix":"422100d6c65e4729","entry":"main","repo":"USTC-StarTeam/RAG-IGBench","repo_kind":"found_in_text","path":"download_images.py","file_url":"https://github.com/USTC-StarTeam/RAG-IGBench/blob/HEAD/download_images.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"422100d6c65e4729"}},{"code_sha256_prefix":"95469ee3b1384512","entry":"url_to_base64","repo":"USTC-StarTeam/RAG-IGBench","repo_kind":"found_in_text","path":"model_generation/claude.py","file_url":"https://github.com/USTC-StarTeam/RAG-IGBench/blob/HEAD/model_generation/claude.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"95469ee3b1384512"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.IR","source":"arxiv_api"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9581,"papers_checked":6264},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}