{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tsa-net-tube-self-attention-network-for","title":"TSA-Net: Tube Self-Attention Network for Action Quality Assessment","arxiv_id":"2201.03746","date":"2022-01-11","proceeding":null,"authors":["Shunli Wang","Dingkang Yang","Peng Zhai","Chixiao Chen","Lihua Zhang"],"abstract":"In recent years, assessing action quality from videos has attracted growing attention in computer vision community and human computer interaction. Most existing approaches usually tackle this problem by directly migrating the model from action recognition tasks, which ignores the intrinsic differences within the feature map such as foreground and background information. To address this issue, we propose a Tube Self-Attention Network (TSA-Net) for action quality assessment (AQA). Specifically, we introduce a single object tracker into AQA and propose the Tube Self-Attention Module (TSA), which can efficiently generate rich spatio-temporal contextual information by adopting sparse feature interactions. The TSA module is embedded in existing video networks to form TSA-Net. Overall, our TSA-Net is with the following merits: 1) High computational efficiency, 2) High flexibility, and 3) The state-of-the art performance. Extensive experiments are conducted on popular action quality assessment datasets including AQA-7 and MTL-AQA. Besides, a dataset named Fall Recognition in Figure Skating (FR-FS) is proposed to explore the basic action assessment in the figure skating scene.","url_abs":"https://arxiv.org/abs/2201.03746v1","url_pdf":"https://arxiv.org/pdf/2201.03746v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tsa-net-tube-self-attention-network-for","repo_url":"https://github.com/Shunli-Wang/TSA-Net","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"tsa-net-tube-self-attention-network-for","repo_url":"https://github.com/MindSpore-scientific-2/code-10/tree/main/TSA_mindspore","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"tsa-net-tube-self-attention-network-for","repo_url":"https://github.com/MindSpore-scientific/code-13/tree/main/TSA_mindspore","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"tsa-net-tube-self-attention-network-for","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/4/TSA_mindspore","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"action-assessment","task_name":"Action Assessment"},{"task_slug":"action-quality-assessment","task_name":"Action Quality Assessment"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"}],"methods":[],"datasets_introduced":[{"slug":"fr-fs","name":"FR-FS","full_name":"Fall Recognition in Figure Skating"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2201.03746","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}