{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/videolstm-convolves-attends-and-flows-for","title":"VideoLSTM Convolves, Attends and Flows for Action Recognition","arxiv_id":"1607.01794","date":"2016-07-06","proceeding":null,"authors":["Zhenyang Li","Efstratios Gavves","Mihir Jain","Cees G. M. Snoek"],"abstract":"We present a new architecture for end-to-end sequence learning of actions in\nvideo, we call VideoLSTM. Rather than adapting the video to the peculiarities\nof established recurrent or convolutional architectures, we adapt the\narchitecture to fit the requirements of the video medium. Starting from the\nsoft-Attention LSTM, VideoLSTM makes three novel contributions. First, video\nhas a spatial layout. To exploit the spatial correlation we hardwire\nconvolutions in the soft-Attention LSTM architecture. Second, motion not only\ninforms us about the action content, but also guides better the attention\ntowards the relevant spatio-temporal locations. We introduce motion-based\nattention. And finally, we demonstrate how the attention from VideoLSTM can be\nused for action localization by relying on just the action class label.\nExperiments and comparisons on challenging datasets for action classification\nand localization support our claims.","url_abs":"http://arxiv.org/abs/1607.01794v1","url_pdf":"http://arxiv.org/pdf/1607.01794v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"videolstm-convolves-attends-and-flows-for","repo_url":"https://github.com/zhenyangli/VideoLSTM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause-Clear"}}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1607.01794","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}