{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/few-shot-action-recognition-via-improved","title":"Few-shot Action Recognition with Permutation-invariant Attention","arxiv_id":"2001.03905","date":"2020-01-12","proceeding":"ECCV 2020 8","authors":["Hongguang Zhang","Li Zhang","Xiaojuan Qi","Hongdong Li","Philip H. S. Torr","Piotr Koniusz"],"abstract":"Many few-shot learning models focus on recognising images. In contrast, we tackle a challenging task of few-shot action recognition from videos. We build on a C3D encoder for spatio-temporal video blocks to capture short-range action patterns. Such encoded blocks are aggregated by permutation-invariant pooling to make our approach robust to varying action lengths and long-range temporal dependencies whose patterns are unlikely to repeat even in clips of the same class. Subsequently, the pooled representations are combined into simple relation descriptors which encode so-called query and support clips. Finally, relation descriptors are fed to the comparator with the goal of similarity learning between query and support clips. Importantly, to re-weight block contributions during pooling, we exploit spatial and temporal attention modules and self-supervision. In naturalistic clips (of the same class) there exists a temporal distribution shift--the locations of discriminative temporal action hotspots vary. Thus, we permute blocks of a clip and align the resulting attention regions with similarly permuted attention regions of non-permuted clip to train the attention mechanism invariant to block (and thus long-term hotspot) permutations. Our method outperforms the state of the art on the HMDB51, UCF101, miniMIT datasets.","url_abs":"https://arxiv.org/abs/2001.03905v3","url_pdf":"https://arxiv.org/pdf/2001.03905v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"few-shot-action-recognition-via-improved","repo_url":"https://github.com/Teddy00888/arn_mindspore","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"few-shot-action-recognition","task_name":"Few Shot Action Recognition"},{"task_slug":"few-shot-learning","task_name":"Few-Shot Learning"},{"task_slug":"few-shot-action-recognition","task_name":"Few-Shot action recognition"},{"task_slug":null,"task_name":"Relation"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/few-shot-action-recognition-on-hmdb51","task":"Few Shot Action Recognition","dataset":"HMDB51","model":"ARN","rank_in_archive_order":7,"of":7,"metrics":{"1:1 Accuracy":"60.6"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-action-recognition-on-kinetics-100","task":"Few Shot Action Recognition","dataset":"Kinetics-100","model":"ARN","rank_in_archive_order":7,"of":8,"metrics":{"Accuracy":"82.4"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-action-recognition-on-ucf101","task":"Few Shot Action Recognition","dataset":"UCF101","model":"ARN","rank_in_archive_order":7,"of":7,"metrics":{"1:1 Accuracy":"83.1"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}