{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/asquery-a-query-based-model-for-action","title":"ASQuery: A Query-based Model for Action Segmentation","arxiv_id":null,"date":"2024-09-30","proceeding":"2024 IEEE International Conference on Multimedia and Expo (ICME) 2024 9","authors":["Gan Ziliang","Jin Lei","Nie Lei","Zheng Wang","Li Zhou","Li Liang","Wang Zhecan","Li Jianshu","Xing Junliang","Zhao Jian"],"abstract":"For the task of temporal action segmentation, existing works commonly treat it as a frame-wise classification problem. In this paper, we propose a straight but effective model namely ASQuery by learning central representation of each action category, which transforms the classification problem to the similarity calculation between category-specific queries and frame features. These central representations are dynamically generated through our Transformer decoder module, endowing them more flexible and comprehensive perception of the whole video. Moreover, we first introduce the boundary query for refining segmentation results, aiding to alleviating the trouble\u0002some over-segmentation problem. ASQuery demonstrates superior performance compared to state-of-the-art models, achieving improvements of 0.9% and 4.1% in the mean metrics on two public action segmentation datasets, i.e., Breakfast and Assembly101, respectively. The source codes are available at https://github.com/zlngan/ASQuery.","url_abs":"https://ieeexplore.ieee.org/document/10687535","url_pdf":"https://ieeexplore.ieee.org/document/10687535","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"asquery-a-query-based-model-for-action","repo_url":"https://github.com/zlngan/ASQuery","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"temporal-action-segmentation","task_name":"Temporal Action Segmentation"},{"task_slug":"model","task_name":"model"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-segmentation-on-assembly101","task":"Action Segmentation","dataset":"Assembly101","model":"ASQuery","rank_in_archive_order":1,"of":7,"metrics":{"Edit":"35.3","F1@10%":"37.8","F1@25%":"35.6","F1@50%":"29.4","MoF":"40.4"},"uses_additional_data":false},{"leaderboard":"/sota/action-segmentation-on-breakfast-1","task":"Action Segmentation","dataset":"Breakfast","model":"ASQuery","rank_in_archive_order":3,"of":37,"metrics":{"Acc":"77.9","Average F1":"74.6","Edit":"78.4","F1@10%":"80.7","F1@25%":"76.5","F1@50%":"66.5"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}