{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lita-language-instructed-temporal","title":"LITA: Language Instructed Temporal-Localization Assistant","arxiv_id":"2403.19046","date":"2024-03-27","proceeding":null,"authors":["De-An Huang","Shijia Liao","Subhashree Radhakrishnan","Hongxu Yin","Pavlo Molchanov","Zhiding Yu","Jan Kautz"],"abstract":"There has been tremendous progress in multimodal Large Language Models (LLMs). Recent works have extended these models to video input with promising instruction following capabilities. However, an important missing piece is temporal localization. These models cannot accurately answer the \"When?\" questions. We identify three key aspects that limit their temporal localization capabilities: (i) time representation, (ii) architecture, and (iii) data. We address these shortcomings by proposing Language Instructed Temporal-Localization Assistant (LITA) with the following features: (1) We introduce time tokens that encode timestamps relative to the video length to better represent time in videos. (2) We introduce SlowFast tokens in the architecture to capture temporal information at fine temporal resolution. (3) We emphasize temporal localization data for LITA. In addition to leveraging existing video datasets with timestamps, we propose a new task, Reasoning Temporal Localization (RTL), along with the dataset, ActivityNet-RTL, for learning and evaluating this task. Reasoning temporal localization requires both the reasoning and temporal localization of Video LLMs. LITA demonstrates strong performance on this challenging task, nearly doubling the temporal mean intersection-over-union (mIoU) of baselines. In addition, we show that our emphasis on temporal localization also substantially improves video-based text generation compared to existing Video LLMs, including a 36% relative improvement of Temporal Understanding. Code is available at: https://github.com/NVlabs/LITA","url_abs":"https://arxiv.org/abs/2403.19046v1","url_pdf":"https://arxiv.org/pdf/2403.19046v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lita-language-instructed-temporal","repo_url":"https://github.com/nvlabs/lita","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"instruction-following","task_name":"Instruction Following"},{"task_slug":"temporal-localization","task_name":"Temporal Localization"},{"task_slug":"text-generation","task_name":"Text Generation"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"video-based-generative-performance","task_name":"Video-based Generative Performance Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-question-answering-on-ovbench","task":"Video Question Answering","dataset":"OVBench","model":"LITA (7B)","rank_in_archive_order":14,"of":16,"metrics":{"AVG":"20.4"},"uses_additional_data":false},{"leaderboard":"/sota/video-based-generative-performance","task":"Video-based Generative Performance Benchmarking","dataset":"VideoInstruct","model":"LITA-13B","rank_in_archive_order":12,"of":23,"metrics":{"Consistency":"3.19","Contextual Understanding":"3.43","Correctness of Information":"2.94","Detail Orientation":"2.98","Temporal Understanding":"2.68","mean":"3.04"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2403.19046","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2403.19046"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/nvlabs/lita","reach":null}],"summary":{"ran_fixture":1,"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"5702eb50c9e0edeb","entry":"iou","repo":"nvlabs/lita","repo_kind":"official","path":"lita/eval/eval_model_rtl.py","file_url":"https://github.com/nvlabs/lita/blob/HEAD/lita/eval/eval_model_rtl.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"5702eb50c9e0edeb"}},{"code_sha256_prefix":"c24fbeafbe82ee96","entry":"parse_start_end_timestamps","repo":"nvlabs/lita","repo_kind":"official","path":"lita/eval/eval_model_rtl.py","file_url":"https://github.com/nvlabs/lita/blob/HEAD/lita/eval/eval_model_rtl.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"c24fbeafbe82ee96"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}