{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/timesuite-improving-mllms-for-long-video","title":"TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning","arxiv_id":"2410.19702","date":"2024-10-25","proceeding":null,"authors":["Xiangyu Zeng","Kunchang Li","Chenting Wang","Xinhao Li","Tianxiang Jiang","Ziang Yan","Songze Li","Yansong Shi","Zhengrong Yue","Yi Wang","Yali Wang","Yu Qiao","LiMin Wang"],"abstract":"Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, a collection of new designs to adapt the existing short-form video MLLMs for long video understanding, including a simple yet efficient framework to process long video sequence, a high-quality video dataset for grounded tuning of MLLMs, and a carefully-designed instruction tuning task to explicitly incorporate the grounding supervision in the traditional QA format. Specifically, based on VideoChat, we propose our long-video MLLM, coined as VideoChat-T, by implementing a token shuffling to compress long video tokens and introducing Temporal Adaptive Position Encoding (TAPE) to enhance the temporal awareness of visual representation. Meanwhile, we introduce the TimePro, a comprehensive grounding-centric instruction tuning dataset composed of 9 tasks and 349k high-quality grounded annotations. Notably, we design a new instruction tuning task type, called Temporal Grounded Caption, to peform detailed video descriptions with the corresponding time stamps prediction. This explicit temporal location prediction will guide MLLM to correctly attend on the visual content when generating description, and thus reduce the hallucination risk caused by the LLMs. Experimental results demonstrate that our TimeSuite provides a successful solution to enhance the long video understanding capability of short-form MLLM, achieving improvement of 5.6% and 6.8% on the benchmarks of Egoschema and VideoMME, respectively. In addition, VideoChat-T exhibits robust zero-shot temporal grounding capabilities, significantly outperforming the existing state-of-the-art MLLMs. After fine-tuning, it performs on par with the traditional supervised expert models.","url_abs":"https://arxiv.org/abs/2410.19702v2","url_pdf":"https://arxiv.org/pdf/2410.19702v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"timesuite-improving-mllms-for-long-video","repo_url":"https://github.com/OpenGVLab/TimeSuite","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":null,"task_name":"EgoSchema"},{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":"highlight-detection","task_name":"Highlight Detection"},{"task_slug":"moment-retrieval","task_name":"Moment Retrieval"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":"zeroshot-video-question-answer","task_name":"Zero-Shot Video Question Answer"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/highlight-detection-on-qvhighlights","task":"Highlight Detection","dataset":"QVHighlights","model":"VideoChat-T (FT)","rank_in_archive_order":19,"of":21,"metrics":{"Hit@1":"55.3","mAP":"27.0"},"uses_additional_data":false},{"leaderboard":"/sota/moment-retrieval-on-charades-sta","task":"Moment Retrieval","dataset":"Charades-STA","model":"VideoChat-T (FT)","rank_in_archive_order":7,"of":25,"metrics":{"R@1 IoU=0.5":"67.1","R@1 IoU=0.7":"43.0"},"uses_additional_data":true},{"leaderboard":"/sota/moment-retrieval-on-charades-sta","task":"Moment Retrieval","dataset":"Charades-STA","model":"VideoChat-T (ZS)","rank_in_archive_order":23,"of":25,"metrics":{"R@1 IoU=0.5":"48.7","R@1 IoU=0.7":"24.0","mIoU":"45.43"},"uses_additional_data":true},{"leaderboard":"/sota/video-question-answering-on-mvbench","task":"Video Question Answering","dataset":"MVBench","model":"VideoChat-T (7B)","rank_in_archive_order":7,"of":22,"metrics":{"Avg.":"59.9"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-question-answer-on-egoschema-1","task":"Zero-Shot Video Question Answer","dataset":"EgoSchema (fullset)","model":"VideoChat-T (7B)","rank_in_archive_order":11,"of":29,"metrics":{"Accuracy":"60.0"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-question-answer-on-egoschema","task":"Zero-Shot Video Question Answer","dataset":"EgoSchema (subset)","model":"VideoChat-T (7B)","rank_in_archive_order":2,"of":14,"metrics":{"Accuracy":"68.4"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-question-answer-on-video-mme-1","task":"Zero-Shot Video Question Answer","dataset":"Video-MME","model":"VideoChat-T (7B)","rank_in_archive_order":11,"of":11,"metrics":{"Accuracy (%)":"55.8"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-question-answer-on-video-mme","task":"Zero-Shot Video Question Answer","dataset":"Video-MME (w/o subs)","model":"VideoChat-T (7B)","rank_in_archive_order":9,"of":9,"metrics":{"Accuracy (%)":"46.3"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.19702","atlas_url":"https://app.syntology.ai/?focus=2410.19702","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.19702"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/OpenGVLab/TimeSuite","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":5,"ran_fixture":1},"by_repo_kind":{"listed":{"samples":6,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a475c4129bb51880","entry":"get_sim","repo":"OpenGVLab/TimeSuite","repo_kind":"listed","path":"models/criterions.py","file_url":"https://github.com/OpenGVLab/TimeSuite/blob/HEAD/models/criterions.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a475c4129bb51880"}},{"code_sha256_prefix":"f64c56c357505158","entry":"interpolate_temporal_pos_embed","repo":"OpenGVLab/TimeSuite","repo_kind":"listed","path":"models/utils.py","file_url":"https://github.com/OpenGVLab/TimeSuite/blob/HEAD/models/utils.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f64c56c357505158"}},{"code_sha256_prefix":"1ffcca73cb98381a","entry":"load_image_from_path","repo":"OpenGVLab/TimeSuite","repo_kind":"listed","path":"dataset/utils.py","file_url":"https://github.com/OpenGVLab/TimeSuite/blob/HEAD/dataset/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1ffcca73cb98381a"}},{"code_sha256_prefix":"5349fd1148712a53","entry":"load_temp_embed_with_mismatch","repo":"OpenGVLab/TimeSuite","repo_kind":"listed","path":"models/utils.py","file_url":"https://github.com/OpenGVLab/TimeSuite/blob/HEAD/models/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5349fd1148712a53"}},{"code_sha256_prefix":"b4d051f7063e9a7b","entry":"pre_text","repo":"OpenGVLab/TimeSuite","repo_kind":"listed","path":"dataset/utils.py","file_url":"https://github.com/OpenGVLab/TimeSuite/blob/HEAD/dataset/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b4d051f7063e9a7b"}},{"code_sha256_prefix":"54d3e29e64d0e87c","entry":"preprocess_para_retrieval_data","repo":"OpenGVLab/TimeSuite","repo_kind":"listed","path":"dataset/pt_dataset.py","file_url":"https://github.com/OpenGVLab/TimeSuite/blob/HEAD/dataset/pt_dataset.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"54d3e29e64d0e87c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}