{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/few-shot-temporal-action-localization-with","title":"Few-Shot Temporal Action Localization with Query Adaptive Transformer","arxiv_id":"2110.10552","date":"2021-10-20","proceeding":null,"authors":["Sauradip Nag","Xiatian Zhu","Tao Xiang"],"abstract":"Existing temporal action localization (TAL) works rely on a large number of training videos with exhaustive segment-level annotation, preventing them from scaling to new classes. As a solution to this problem, few-shot TAL (FS-TAL) aims to adapt a model to a new class represented by as few as a single video. Exiting FS-TAL methods assume trimmed training videos for new classes. However, this setting is not only unnatural actions are typically captured in untrimmed videos, but also ignores background video segments containing vital contextual cues for foreground action segmentation. In this work, we first propose a new FS-TAL setting by proposing to use untrimmed training videos. Further, a novel FS-TAL model is proposed which maximizes the knowledge transfer from training classes whilst enabling the model to be dynamically adapted to both the new class and each video of that class simultaneously. This is achieved by introducing a query adaptive Transformer in the model. Extensive experiments on two action localization benchmarks demonstrate that our method can outperform all the state of the art alternatives significantly in both single-domain and cross-domain scenarios. The source code can be found in https://github.com/sauradip/fewshotQAT","url_abs":"https://arxiv.org/abs/2110.10552v1","url_pdf":"https://arxiv.org/pdf/2110.10552v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"few-shot-temporal-action-localization-with","repo_url":"https://github.com/sauradip/fewshotQAT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"few-shot-temporal-action-localization","task_name":"Few Shot Temporal Action Localization"},{"task_slug":"fine-grained-action-detection","task_name":"Fine-Grained Action Detection"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"temporal-localization","task_name":"Temporal Localization"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/few-shot-temporal-action-localization-on","task":"Few Shot Temporal Action Localization","dataset":"ActivityNet","model":"FS-QAT","rank_in_archive_order":1,"of":1,"metrics":{"mIoU":"38.5"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-temporal-action-localization-on-1","task":"Few Shot Temporal Action Localization","dataset":"THUMOS14","model":"FS-QAT","rank_in_archive_order":1,"of":1,"metrics":{"mIoU":"30.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2110.10552","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}