{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/collaborative-spatio-temporal-feature","title":"Collaborative Spatio-temporal Feature Learning for Video Action Recognition","arxiv_id":"1903.01197","date":"2019-03-04","proceeding":null,"authors":["Chao Li","Qiaoyong Zhong","Di Xie","ShiLiang Pu"],"abstract":"Spatio-temporal feature learning is of central importance for action\nrecognition in videos. Existing deep neural network models either learn spatial\nand temporal features independently (C2D) or jointly with unconstrained\nparameters (C3D). In this paper, we propose a novel neural operation which\nencodes spatio-temporal features collaboratively by imposing a weight-sharing\nconstraint on the learnable parameters. In particular, we perform 2D\nconvolution along three orthogonal views of volumetric video data,which learns\nspatial appearance and temporal motion cues respectively. By sharing the\nconvolution kernels of different views, spatial and temporal features are\ncollaboratively learned and thus benefit from each other. The complementary\nfeatures are subsequently fused by a weighted summation whose coefficients are\nlearned end-to-end. Our approach achieves state-of-the-art performance on\nlarge-scale benchmarks and won the 1st place in the Moments in Time Challenge\n2018. Moreover, based on the learned coefficients of different views, we are\nable to quantify the contributions of spatial and temporal features. This\nanalysis sheds light on interpretability of the model and may also guide the\nfuture design of algorithm for video recognition.","url_abs":"http://arxiv.org/abs/1903.01197v1","url_pdf":"http://arxiv.org/pdf/1903.01197v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"collaborative-spatio-temporal-feature","repo_url":"https://github.com/hikvision-research/cost","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-recognition","task_name":"Video Recognition"}],"methods":[{"method_slug":"interpretability","method_name":"Interpretability"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.01197","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}