{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spatio-temporal-relation-modeling-for-few","title":"Spatio-temporal Relation Modeling for Few-shot Action Recognition","arxiv_id":"2112.05132","date":"2021-12-09","proceeding":"CVPR 2022 1","authors":["Anirudh Thatipelli","Sanath Narayan","Salman Khan","Rao Muhammad Anwer","Fahad Shahbaz Khan","Bernard Ghanem"],"abstract":"We propose a novel few-shot action recognition framework, STRM, which enhances class-specific feature discriminability while simultaneously learning higher-order temporal representations. The focus of our approach is a novel spatio-temporal enrichment module that aggregates spatial and temporal contexts with dedicated local patch-level and global frame-level feature enrichment sub-modules. Local patch-level enrichment captures the appearance-based characteristics of actions. On the other hand, global frame-level enrichment explicitly encodes the broad temporal context, thereby capturing the relevant object features over time. The resulting spatio-temporally enriched representations are then utilized to learn the relational matching between query and support action sub-sequences. We further introduce a query-class similarity classifier on the patch-level enriched features to enhance class-specific feature discriminability by reinforcing the feature learning at different stages in the proposed framework. Experiments are performed on four few-shot action recognition benchmarks: Kinetics, SSv2, HMDB51 and UCF101. Our extensive ablation study reveals the benefits of the proposed contributions. Furthermore, our approach sets a new state-of-the-art on all four benchmarks. On the challenging SSv2 benchmark, our approach achieves an absolute gain of $3.5\\%$ in classification accuracy, as compared to the best existing method in the literature. Our code and models are available at https://github.com/Anirudh257/strm.","url_abs":"https://arxiv.org/abs/2112.05132v2","url_pdf":"https://arxiv.org/pdf/2112.05132v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spatio-temporal-relation-modeling-for-few","repo_url":"https://github.com/Anirudh257/strm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"few-shot-action-recognition","task_name":"Few Shot Action Recognition"},{"task_slug":"few-shot-action-recognition","task_name":"Few-Shot action recognition"},{"task_slug":null,"task_name":"Relation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/few-shot-action-recognition-on-hmdb51","task":"Few Shot Action Recognition","dataset":"HMDB51","model":"STRM","rank_in_archive_order":1,"of":7,"metrics":{"1:1 Accuracy":"77.3"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-action-recognition-on-kinetics-100","task":"Few Shot Action Recognition","dataset":"Kinetics-100","model":"STRM","rank_in_archive_order":3,"of":8,"metrics":{"Accuracy":"86.7"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-action-recognition-on-something","task":"Few Shot Action Recognition","dataset":"Something-Something-100","model":"STRM","rank_in_archive_order":2,"of":5,"metrics":{"1:1 Accuracy":"68.1"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-action-recognition-on-ucf101","task":"Few Shot Action Recognition","dataset":"UCF101","model":"STRM","rank_in_archive_order":1,"of":7,"metrics":{"1:1 Accuracy":"96.8"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2112.05132","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}