{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spatiotemporal-residual-networks-for-video","title":"Spatiotemporal Residual Networks for Video Action Recognition","arxiv_id":"1611.02155","date":"2016-11-07","proceeding":"NeurIPS 2016 12","authors":["Christoph Feichtenhofer","Axel Pinz","Richard P. Wildes"],"abstract":"Two-stream Convolutional Networks (ConvNets) have shown strong performance\nfor human action recognition in videos. Recently, Residual Networks (ResNets)\nhave arisen as a new technique to train extremely deep architectures. In this\npaper, we introduce spatiotemporal ResNets as a combination of these two\napproaches. Our novel architecture generalizes ResNets for the spatiotemporal\ndomain by introducing residual connections in two ways. First, we inject\nresidual connections between the appearance and motion pathways of a two-stream\narchitecture to allow spatiotemporal interaction between the two streams.\nSecond, we transform pretrained image ConvNets into spatiotemporal networks by\nequipping these with learnable convolutional filters that are initialized as\ntemporal residual connections and operate on adjacent feature maps in time.\nThis approach slowly increases the spatiotemporal receptive field as the depth\nof the model increases and naturally integrates image ConvNet design\nprinciples. The whole model is trained end-to-end to allow hierarchical\nlearning of complex spatiotemporal features. We evaluate our novel\nspatiotemporal ResNet using two widely used action recognition benchmarks where\nit exceeds the previous state-of-the-art.","url_abs":"http://arxiv.org/abs/1611.02155v1","url_pdf":"http://arxiv.org/pdf/1611.02155v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spatiotemporal-residual-networks-for-video","repo_url":"https://github.com/feichtenhofer/st-resnet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-hmdb-51","task":"Action Recognition","dataset":"HMDB-51","model":"ST-ResNet + IDT","rank_in_archive_order":54,"of":77,"metrics":{"Average accuracy of 3 splits":"70.3"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"ST-ResNet + IDT","rank_in_archive_order":51,"of":91,"metrics":{"3-fold Accuracy":"94.6"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1611.02155","atlas_url":"https://app.syntology.ai/?focus=1611.02155","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}