{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/convolutional-two-stream-network-fusion-for","title":"Convolutional Two-Stream Network Fusion for Video Action Recognition","arxiv_id":"1604.06573","date":"2016-04-22","proceeding":"CVPR 2016 6","authors":["Christoph Feichtenhofer","Axel Pinz","Andrew Zisserman"],"abstract":"Recent applications of Convolutional Neural Networks (ConvNets) for human\naction recognition in videos have proposed different solutions for\nincorporating the appearance and motion information. We study a number of ways\nof fusing ConvNet towers both spatially and temporally in order to best take\nadvantage of this spatio-temporal information. We make the following findings:\n(i) that rather than fusing at the softmax layer, a spatial and temporal\nnetwork can be fused at a convolution layer without loss of performance, but\nwith a substantial saving in parameters; (ii) that it is better to fuse such\nnetworks spatially at the last convolutional layer than earlier, and that\nadditionally fusing at the class prediction layer can boost accuracy; finally\n(iii) that pooling of abstract convolutional features over spatiotemporal\nneighbourhoods further boosts performance. Based on these studies we propose a\nnew ConvNet architecture for spatiotemporal fusion of video snippets, and\nevaluate its performance on standard benchmarks where this architecture\nachieves state-of-the-art results.","url_abs":"http://arxiv.org/abs/1604.06573v2","url_pdf":"http://arxiv.org/pdf/1604.06573v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"convolutional-two-stream-network-fusion-for","repo_url":"https://github.com/feichtenhofer/twostreamfusion","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"convolutional-two-stream-network-fusion-for","repo_url":"https://github.com/tomar840/two-stream-fusion-for-action-recognition-in-videos","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"two","task_name":"Vocal Bursts Valence Prediction"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-hmdb-51","task":"Action Recognition","dataset":"HMDB-51","model":"S:VGG-16, T:VGG-16 (ImageNet pretrained)","rank_in_archive_order":62,"of":77,"metrics":{"Average accuracy of 3 splits":"65.4"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"S:VGG-16, T:VGG-16 (ImageNet pretrain)","rank_in_archive_order":63,"of":91,"metrics":{"3-fold Accuracy":"92.5"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1604.06573","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}