{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-latent-sub-events-in-activity-videos","title":"Learning Latent Sub-events in Activity Videos Using Temporal Attention Filters","arxiv_id":"1605.08140","date":"2016-05-26","proceeding":null,"authors":["AJ Piergiovanni","Chenyou Fan","Michael S. Ryoo"],"abstract":"In this paper, we newly introduce the concept of temporal attention filters,\nand describe how they can be used for human activity recognition from videos.\nMany high-level activities are often composed of multiple temporal parts (e.g.,\nsub-events) with different duration/speed, and our objective is to make the\nmodel explicitly learn such temporal structure using multiple attention filters\nand benefit from them. Our temporal filters are designed to be fully\ndifferentiable, allowing end-of-end training of the temporal filters together\nwith the underlying frame-based or segment-based convolutional neural network\narchitectures. This paper presents an approach of learning a set of optimal\nstatic temporal attention filters to be shared across different videos, and\nextends this approach to dynamically adjust attention filters per testing video\nusing recurrent long short-term memory networks (LSTMs). This allows our\ntemporal attention filters to learn latent sub-events specific to each\nactivity. We experimentally confirm that the proposed concept of temporal\nattention filters benefits the activity recognition, and we visualize the\nlearned latent sub-events.","url_abs":"http://arxiv.org/abs/1605.08140v3","url_pdf":"http://arxiv.org/pdf/1605.08140v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-latent-sub-events-in-activity-videos","repo_url":"https://github.com/piergiaj/latent-subevents","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"activity-recognition-in-videos","task_name":"Activity Recognition In Videos"},{"task_slug":"human-activity-recognition","task_name":"Human Activity Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/activity-recognition-in-videos-on-dogcentric","task":"Activity Recognition In Videos","dataset":"DogCentric","model":"VTFSA","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"98.55"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}