{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learnable-pooling-methods-for-video","title":"Learnable Pooling Methods for Video Classification","arxiv_id":"1810.00530","date":"2018-10-01","proceeding":null,"authors":["Sebastian Kmiec","Juhan Bae","Ruijian An"],"abstract":"We introduce modifications to state-of-the-art approaches to aggregating\nlocal video descriptors by using attention mechanisms and function\napproximations. Rather than using ensembles of existing architectures, we\nprovide an insight on creating new architectures. We demonstrate our solutions\nin the \"The 2nd YouTube-8M Video Understanding Challenge\", by using frame-level\nvideo and audio descriptors. We obtain testing accuracy similar to the state of\nthe art, while meeting budget constraints, and touch upon strategies to improve\nthe state of the art. Model implementations are available in\nhttps://github.com/pomonam/LearnablePoolingMethods.","url_abs":"http://arxiv.org/abs/1810.00530v1","url_pdf":"http://arxiv.org/pdf/1810.00530v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learnable-pooling-methods-for-video","repo_url":"https://github.com/pomonam/LearnablePoolingMethods","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"video-classification","task_name":"Video Classification"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}