{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-gating-convnet-for-two-stream-based","title":"Learning Gating ConvNet for Two-Stream based Methods in Action Recognition","arxiv_id":"1709.03655","date":"2017-09-12","proceeding":null,"authors":["Jiagang Zhu","Wei Zou","Zheng Zhu"],"abstract":"For the two-stream style methods in action recognition, fusing the two\nstreams' predictions is always by the weighted averaging scheme. This fusion\nmethod with fixed weights lacks of pertinence to different action videos and\nalways needs trial and error on the validation set. In order to enhance the\nadaptability of two-stream ConvNets and improve its performance, an end-to-end\ntrainable gated fusion method, namely gating ConvNet, for the two-stream\nConvNets is proposed in this paper based on the MoE (Mixture of Experts)\ntheory. The gating ConvNet takes the combination of feature maps from the same\nlayer of the spatial and the temporal nets as input and adopts ReLU (Rectified\nLinear Unit) as the gating output activation function. To reduce the\nover-fitting of gating ConvNet caused by the redundancy of parameters, a new\nmulti-task learning method is designed, which jointly learns the gating fusion\nweights for the two streams and learns the gating ConvNet for action\nclassification. With our gated fusion method and multi-task learning approach,\na high accuracy of 94.5% is achieved on the dataset UCF101.","url_abs":"http://arxiv.org/abs/1709.03655v2","url_pdf":"http://arxiv.org/pdf/1709.03655v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-gating-convnet-for-two-stream-based","repo_url":"https://github.com/zhujiagang/gating-ConvNet-code","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"mixture-of-experts","task_name":"Mixture-of-Experts"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"two","task_name":"Vocal Bursts Valence Prediction"}],"methods":[{"method_slug":"relu","method_name":"ReLU"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}