{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learnable-pooling-with-context-gating-for","title":"Learnable pooling with Context Gating for video classification","arxiv_id":"1706.06905","date":"2017-06-21","proceeding":null,"authors":["Antoine Miech","Ivan Laptev","Josef Sivic"],"abstract":"Current methods for video analysis often extract frame-level features using\npre-trained convolutional neural networks (CNNs). Such features are then\naggregated over time e.g., by simple temporal averaging or more sophisticated\nrecurrent neural networks such as long short-term memory (LSTM) or gated\nrecurrent units (GRU). In this work we revise existing video representations\nand study alternative methods for temporal aggregation. We first explore\nclustering-based aggregation layers and propose a two-stream architecture\naggregating audio and visual features. We then introduce a learnable non-linear\nunit, named Context Gating, aiming to model interdependencies among network\nactivations. Our experimental results show the advantage of both improvements\nfor the task of video classification. In particular, we evaluate our method on\nthe large-scale multi-modal Youtube-8M v2 dataset and outperform all other\nmethods in the Youtube 8M Large-Scale Video Understanding challenge.","url_abs":"http://arxiv.org/abs/1706.06905v2","url_pdf":"http://arxiv.org/pdf/1706.06905v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learnable-pooling-with-context-gating-for","repo_url":"https://github.com/antoine77340/Youtube-8M-WILLOW","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"learnable-pooling-with-context-gating-for","repo_url":"https://github.com/antoine77340/LOUPE","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"learnable-pooling-with-context-gating-for","repo_url":"https://github.com/mogadmd/YT8M","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"learnable-pooling-with-context-gating-for","repo_url":"https://github.com/pomonam/AttentionCluster","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"learnable-pooling-with-context-gating-for","repo_url":"https://github.com/MindSpore-paper-code-3/code1/tree/main/AttentionCluster","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"video-classification","task_name":"Video Classification"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.06905","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}