{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rtv-bench-benchmarking-mllm-continuous","title":"RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video","arxiv_id":"2505.02064","date":"2025-05-04","proceeding":null,"authors":["Shuhang Xun","Sicheng Tao","Jungang Li","Yibo Shi","Zhixin Lin","Zhanhui Zhu","Yibo Yan","Hanqian Li","Linghao Zhang","Shikang Wang","Yixin Liu","Hanbo Zhang","Ying Ma","Xuming Hu"],"abstract":"Multimodal Large Language Models (MLLMs) increasingly excel at perception, understanding, and reasoning. However, current benchmarks inadequately evaluate their ability to perform these tasks continuously in dynamic, real-world environments. To bridge this gap, we introduce RTV-Bench, a fine-grained benchmark for MLLM real-time video analysis. RTV-Bench uses three key principles: (1) Multi-Timestamp Question Answering (MTQA), where answers evolve with scene changes; (2) Hierarchical Question Structure, combining basic and advanced queries; and (3) Multi-dimensional Evaluation, assessing the ability of continuous perception, understanding, and reasoning. RTV-Bench contains 552 diverse videos (167.2 hours) and 4,631 high-quality QA pairs. We evaluated leading MLLMs, including proprietary (GPT-4o, Gemini 2.0), open-source offline (Qwen2.5-VL, VideoLLaMA3), and open-source real-time (VITA-1.5, InternLM-XComposer2.5-OmniLive) models. Experiment results show open-source real-time models largely outperform offline ones but still trail top proprietary models. Our analysis also reveals that larger model size or higher frame sampling rates do not significantly boost RTV-Bench performance, sometimes causing slight decreases. This underscores the need for better model architectures optimized for video stream processing and long sequences to advance real-time video analysis with MLLMs. Our benchmark toolkit is available at: https://github.com/LJungang/RTV-Bench.","url_abs":"https://arxiv.org/abs/2505.02064v2","url_pdf":"https://arxiv.org/pdf/2505.02064v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rtv-bench-benchmarking-mllm-continuous","repo_url":"https://github.com/ljungang/rtv-bench","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2505.02064","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.02064"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ljungang/rtv-bench","reach":null}],"summary":{"ran_draft_wrong":3,"ran_violates":1,"ran_honours":1},"by_repo_kind":{"official":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"67944deda32bd0dc","entry":"ensure_dir","repo":"ljungang/rtv-bench","repo_kind":"official","path":"scripts/eval/compute_score.py","file_url":"https://github.com/ljungang/rtv-bench/blob/HEAD/scripts/eval/compute_score.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"67944deda32bd0dc"}},{"code_sha256_prefix":"b0d53ecfb33f3777","entry":"expand_inputs","repo":"ljungang/rtv-bench","repo_kind":"official","path":"scripts/eval/compute_acc.py","file_url":"https://github.com/ljungang/rtv-bench/blob/HEAD/scripts/eval/compute_acc.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b0d53ecfb33f3777"}},{"code_sha256_prefix":"61c52492c49a4f53","entry":"expand_inputs","repo":"ljungang/rtv-bench","repo_kind":"official","path":"scripts/eval/compute_score.py","file_url":"https://github.com/ljungang/rtv-bench/blob/HEAD/scripts/eval/compute_score.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"61c52492c49a4f53"}},{"code_sha256_prefix":"38691cddc327b5f8","entry":"macro_avg","repo":"ljungang/rtv-bench","repo_kind":"official","path":"scripts/eval/compute_acc.py","file_url":"https://github.com/ljungang/rtv-bench/blob/HEAD/scripts/eval/compute_acc.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"38691cddc327b5f8"}},{"code_sha256_prefix":"7e6c2b4cc1362b87","entry":"safe_rate","repo":"ljungang/rtv-bench","repo_kind":"official","path":"scripts/eval/compute_acc.py","file_url":"https://github.com/ljungang/rtv-bench/blob/HEAD/scripts/eval/compute_acc.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7e6c2b4cc1362b87"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}