{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-prompt-multi-modal-tracking","title":"Visual Prompt Multi-Modal Tracking","arxiv_id":"2303.10826","date":"2023-03-20","proceeding":"CVPR 2023 1","authors":["Jiawen Zhu","Simiao Lai","Xin Chen","Dong Wang","Huchuan Lu"],"abstract":"Visible-modal object tracking gives rise to a series of downstream multi-modal tracking tributaries. To inherit the powerful representations of the foundation model, a natural modus operandi for multi-modal tracking is full fine-tuning on the RGB-based parameters. Albeit effective, this manner is not optimal due to the scarcity of downstream data and poor transferability, etc. In this paper, inspired by the recent success of the prompt learning in language models, we develop Visual Prompt multi-modal Tracking (ViPT), which learns the modal-relevant prompts to adapt the frozen pre-trained foundation model to various downstream multimodal tracking tasks. ViPT finds a better way to stimulate the knowledge of the RGB-based model that is pre-trained at scale, meanwhile only introducing a few trainable parameters (less than 1% of model parameters). ViPT outperforms the full fine-tuning paradigm on multiple downstream tracking tasks including RGB+Depth, RGB+Thermal, and RGB+Event tracking. Extensive experiments show the potential of visual prompt learning for multi-modal tracking, and ViPT can achieve state-of-the-art performance while satisfying parameter efficiency. Code and models are available at https://github.com/jiawen-zhu/ViPT.","url_abs":"https://arxiv.org/abs/2303.10826v2","url_pdf":"https://arxiv.org/pdf/2303.10826v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-prompt-multi-modal-tracking","repo_url":"https://github.com/jiawen-zhu/vipt","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object-tracking","task_name":"Object Tracking"},{"task_slug":"prompt-learning","task_name":"Prompt Learning"},{"task_slug":"rgb-t-tracking","task_name":"Rgb-T Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/rgb-t-tracking-on-lasher","task":"Rgb-T Tracking","dataset":"LasHeR","model":"ViPT","rank_in_archive_order":32,"of":39,"metrics":{"Precision":"65.1","Success":"52.5"},"uses_additional_data":false},{"leaderboard":"/sota/rgb-t-tracking-on-rgbt234","task":"Rgb-T Tracking","dataset":"RGBT234","model":"ViPT","rank_in_archive_order":33,"of":42,"metrics":{"Precision":"83.5","Success":"61.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2303.10826","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2303.10826"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/jiawen-zhu/ViPT","reach":null}],"summary":{"ran":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"dca6fc3ac97bba2e","entry":"ViPTrack","repo":"jiawen-zhu/ViPT","repo_kind":"official","path":"lib/models/vipt/ostrack_prompt.py","file_url":"https://github.com/jiawen-zhu/ViPT/blob/HEAD/lib/models/vipt/ostrack_prompt.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"dca6fc3ac97bba2e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}