{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/side4video-spatial-temporal-side-network-for","title":"Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning","arxiv_id":"2311.15769","date":"2023-11-27","proceeding":null,"authors":["Huanjin Yao","Wenhao Wu","Zhiheng Li"],"abstract":"Large pre-trained vision models achieve impressive success in computer vision. However, fully fine-tuning large models for downstream tasks, particularly in video understanding, can be prohibitively computationally expensive. Recent studies turn their focus towards efficient image-to-video transfer learning. Nevertheless, existing efficient fine-tuning methods lack attention to training memory usage and exploration of transferring a larger model to the video domain. In this paper, we present a novel Spatial-Temporal Side Network for memory-efficient fine-tuning large image models to video understanding, named Side4Video. Specifically, we introduce a lightweight spatial-temporal side network attached to the frozen vision model, which avoids the backpropagation through the heavy pre-trained model and utilizes multi-level spatial features from the original image model. Extremely memory-efficient architecture enables our method to reduce 75% memory usage than previous adapter-based methods. In this way, we can transfer a huge ViT-E (4.4B) for video understanding tasks which is 14x larger than ViT-L (304M). Our approach achieves remarkable performance on various video datasets across unimodal and cross-modal tasks (i.e., action recognition and text-video retrieval), especially in Something-Something V1&V2 (67.3% & 74.6%), Kinetics-400 (88.6%), MSR-VTT (52.3%), MSVD (56.1%) and VATEX (68.8%). We release our code at https://github.com/HJYao00/Side4Video.","url_abs":"https://arxiv.org/abs/2311.15769v1","url_pdf":"https://arxiv.org/pdf/2311.15769v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"side4video-spatial-temporal-side-network-for","repo_url":"https://github.com/HJYao00/Side4Video","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"side4video-spatial-temporal-side-network-for","repo_url":"https://github.com/whwu95/ATM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"Side4Video (EVA, ViT-E/14)","rank_in_archive_order":23,"of":207,"metrics":{"Acc@1":"88.6","Acc@5":"98.2"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"Side4Video (EVA ViT-E/14","rank_in_archive_order":3,"of":74,"metrics":{"Top 1 Accuracy":"67.3","Top 5 Accuracy":"88.8"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"Side4Video (EVA ViT-E/14)","rank_in_archive_order":10,"of":123,"metrics":{"Top-1 Accuracy":"75.2","Top-5 Accuracy":"94.0"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-msr-vtt-1ka","task":"Video Retrieval","dataset":"MSR-VTT-1kA","model":"Side4Video","rank_in_archive_order":14,"of":63,"metrics":{"text-to-video Mean Rank":"12.8","text-to-video Median Rank":"1.0","text-to-video R@1":"52.3","text-to-video R@10":"84.2","text-to-video R@5":"75.5"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-msvd","task":"Video Retrieval","dataset":"MSVD","model":"Side4Video","rank_in_archive_order":8,"of":24,"metrics":{"text-to-video Mean Rank":"8.4","text-to-video Median Rank":"1.0","text-to-video R@1":"56.1","text-to-video R@10":"88.8","text-to-video R@5":"81.7"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-vatex","task":"Video Retrieval","dataset":"VATEX","model":"Side4Video","rank_in_archive_order":7,"of":13,"metrics":{"text-to-video MedianR":"2.7","text-to-video R@1":"68.8","text-to-video R@10":"97.0","text-to-video R@5":"93.5","text-to-video R@50":"1.0"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2311.15769","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}