{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/locate-gat-modeling-multi-scale-local-context","title":"LoCATe-GAT: Modeling Multi-Scale Local Context and Action Relationships for Zero-Shot Action Recognition","arxiv_id":null,"date":"2024-11-27","proceeding":"IEEE Transactions on Emerging Topics in Computational Intelligence (TETCI) 2024 11","authors":["Sandipan Sarma","Divyam Singal","Arijit Sur"],"abstract":"The increasing number of actions in the real world makes it difficult for traditional deep-learning models to recognize unseen actions. Recently, pretrained contrastive image-based visual-language (I-VL) models have been adapted for efficient “zero-shot” scene understanding. Pairing such models with transformers to implement temporal modeling has been rewarding for zero-shot action recognition (ZSAR). However, the significance of modeling the local spatial context of objects and action environments remains unexplored. In this work, we propose a ZSAR framework called LoCATe-GAT, comprising a novel Local Context-Aggregating Temporal transformer (LoCATe) and a Graph Attention Network (GAT). Specifically, image and text encodings extracted from a pretrained I-VL model are used as inputs for LoCATe-GAT. Motivated by the observation that object-centric and environmental contexts drive both distinguishability and functional similarity between actions, LoCATe captures multi-scale local context using dilated convolutional layers during temporal modeling. Furthermore, the proposed GAT models semantic relationships between classes and achieves a strong synergy with the video embeddings produced by LoCATe. Extensive experiments on four widely-used benchmarks – UCF101, HMDB51, ActivityNet, and Kinetics – show we achieve state-of-the-art results. Specifically, we obtain relative gains of 3.8% and 4.8% on these datasets in conventional and 16.6% on UCF101in generalized ZSAR settings. For large-scale datasets like ActivityNet and Kinetics, our method achieves a relative gain of 31.8% and 27.9%, respectively, over the previous methods. Additionally, we gain 25.3% and 18.4%\r\non UCF101 and HMDB51 as per the recent “TruZe” evaluation protocol.","url_abs":"https://ieeexplore.ieee.org/document/10769605","url_pdf":"https://ieeexplore.ieee.org/document/10769605","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"locate-gat-modeling-multi-scale-local-context","repo_url":"https://github.com/sandipan211/LoCATe-GAT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"graph-attention","task_name":"Graph Attention"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"zero-shot-action-recognition","task_name":"Zero-Shot Action Recognition"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"gat","method_name":"GAT"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/zero-shot-action-recognition-on-activitynet","task":"Zero-Shot Action Recognition","dataset":"ActivityNet","model":"LoCATe-GAT","rank_in_archive_order":3,"of":5,"metrics":{"Top-1 Accuracy":"73.8"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-action-recognition-on-hmdb51","task":"Zero-Shot Action Recognition","dataset":"HMDB51","model":"LoCATe-GAT","rank_in_archive_order":12,"of":29,"metrics":{"Top-1 Accuracy":"50.7"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-action-recognition-on-kinetics","task":"Zero-Shot Action Recognition","dataset":"Kinetics","model":"LoCATe-GAT","rank_in_archive_order":11,"of":20,"metrics":{"Top-1 Accuracy":"58.7"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-action-recognition-on-ucf101","task":"Zero-Shot Action Recognition","dataset":"UCF101","model":"LoCATe-GAT","rank_in_archive_order":13,"of":35,"metrics":{"Top-1 Accuracy":"76.0"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}