{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/modeling-spatial-temporal-clues-in-a-hybrid","title":"Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification","arxiv_id":"1504.01561","date":"2015-04-07","proceeding":null,"authors":["Zuxuan Wu","Xi Wang","Yu-Gang Jiang","Hao Ye","xiangyang xue"],"abstract":"Classifying videos according to content semantics is an important problem\nwith a wide range of applications. In this paper, we propose a hybrid deep\nlearning framework for video classification, which is able to model static\nspatial information, short-term motion, as well as long-term temporal clues in\nthe videos. Specifically, the spatial and the short-term motion features are\nextracted separately by two Convolutional Neural Networks (CNN). These two\ntypes of CNN-based features are then combined in a regularized feature fusion\nnetwork for classification, which is able to learn and utilize feature\nrelationships for improved performance. In addition, Long Short Term Memory\n(LSTM) networks are applied on top of the two features to further model\nlonger-term temporal clues. The main contribution of this work is the hybrid\nlearning framework that can model several important aspects of the video data.\nWe also show that (1) combining the spatial and the short-term motion features\nin the regularized fusion network is better than direct classification and\nfusion using the CNN with a softmax layer, and (2) the sequence-based LSTM is\nhighly complementary to the traditional classification strategy without\nconsidering the temporal frame orders. Extensive experiments are conducted on\ntwo popular and challenging benchmarks, the UCF-101 Human Actions and the\nColumbia Consumer Videos (CCV). On both benchmarks, our framework achieves\nto-date the best reported performance: $91.3\\%$ on the UCF-101 and $83.5\\%$ on\nthe CCV.","url_abs":"http://arxiv.org/abs/1504.01561v1","url_pdf":"http://arxiv.org/pdf/1504.01561v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"modeling-spatial-temporal-clues-in-a-hybrid","repo_url":"https://github.com/tejgvsl/Camera-motion-classification-in-a-video-file","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1504.01561","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}