{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vispeak-visual-instruction-feedback-in","title":"ViSpeak: Visual Instruction Feedback in Streaming Videos","arxiv_id":"2503.12769","date":"2025-03-17","proceeding":null,"authors":["Shenghao Fu","Qize Yang","Yuan-Ming Li","Yi-Xing Peng","Kun-Yu Lin","Xihan Wei","Jian-Fang Hu","Xiaohua Xie","Wei-Shi Zheng"],"abstract":"Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming video understanding from a new perspective and propose a novel task named Visual Instruction Feedback in which models should be aware of visual contents and learn to extract instructions from them. For example, when users wave their hands to agents, agents should recognize the gesture and start conversations with welcome information. Thus, following instructions in visual modality greatly enhances user-agent interactions. To facilitate research, we define seven key subtasks highly relevant to visual modality and collect the ViSpeak-Instruct dataset for training and the ViSpeak-Bench for evaluation. Further, we propose the ViSpeak model, which is a SOTA streaming video understanding LMM with GPT-4o-level performance on various streaming video understanding benchmarks. After finetuning on our ViSpeak-Instruct dataset, ViSpeak is equipped with basic visual instruction feedback ability, serving as a solid baseline for future research.","url_abs":"https://arxiv.org/abs/2503.12769v1","url_pdf":"https://arxiv.org/pdf/2503.12769v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vispeak-visual-instruction-feedback-in","repo_url":"https://github.com/thunlp-mt/streamingbench","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"streaming-video-understanding","task_name":"Streaming video understanding"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[{"method_slug":"aware","method_name":"AWARE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2503.12769","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.12769"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/thunlp-mt/streamingbench","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":1,"ran_draft_wrong":1,"ran_fixture":1,"unverified":1},"by_repo_kind":{"listed":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"57e95166ade5f25c","entry":"build_transform","repo":"thunlp-mt/streamingbench","repo_kind":"listed","path":"src/model/InternVL.py","file_url":"https://github.com/thunlp-mt/streamingbench/blob/HEAD/src/model/InternVL.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"57e95166ade5f25c"}},{"code_sha256_prefix":"d610b5eabe0c9db1","entry":"dynamic_preprocess","repo":"thunlp-mt/streamingbench","repo_kind":"listed","path":"src/model/InternVL.py","file_url":"https://github.com/thunlp-mt/streamingbench/blob/HEAD/src/model/InternVL.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d610b5eabe0c9db1"}},{"code_sha256_prefix":"bd77f5f8067f18e9","entry":"find_closest_aspect_ratio","repo":"thunlp-mt/streamingbench","repo_kind":"listed","path":"src/model/InternVL.py","file_url":"https://github.com/thunlp-mt/streamingbench/blob/HEAD/src/model/InternVL.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bd77f5f8067f18e9"}},{"code_sha256_prefix":"6a925b5c04f64ad4","entry":"Kangaroo_Run","repo":"thunlp-mt/streamingbench","repo_kind":"listed","path":"src/model/Kangaroo.py","file_url":"https://github.com/thunlp-mt/streamingbench/blob/HEAD/src/model/Kangaroo.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6a925b5c04f64ad4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}