{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ae-net-adjoint-enhancement-network-for","title":"AE-Net:Adjoint Enhancement Network for Efficient Action Recognition in Video Understanding","arxiv_id":null,"date":"2022-07-21","proceeding":"TMM 2022 7","authors":["Bin Wang","Chunsheng Liu","Faliang Chang","Wenqian Wang and Nanjun Li"],"abstract":"Action recognition in video understanding is a challenging task, largely because of the complexity and difficulty in\r\ntemporal modeling, making it suffer from motion information loss\r\nand misalignment of temporal attention in spatial dimensions. To\r\novercome these difficulties, we propose a novel temporal modeling\r\nmethod called Adjoint Enhancement Network (AE-Net), which\r\ncan fully explore clues of motion and time in the long-range\r\nstructure. The AE-Net mainly consists of two new modules:\r\nthe Initial Adjoint Enhancement Module (IAE-Module), which\r\ndeals with shallow features; and the Global Adjoint Enhancement\r\nModule (GAE-Module), which deals with global features. With\r\na novel mechanism of parallel spatio-temporal convolution and\r\ndifference fusion, the IAE-Module is to enhance the degree of\r\nmotion transformation in shallow network features, exciting the\r\npotential of motion flow and avoiding motion information loss.\r\nThe GAE-Module is proposed to improve the local temporal\r\nrepresentation in long-range structures by feeding the enhanced\r\nfeature differences into a spatial cascade module with residuals\r\nto resolve the misalignment of temporal attention in the spatial\r\ndimension.The experimental results show that our AE-Net can\r\nachieve state-of-the-art results in Something-Something V1, UCF101 and HMDB-51 datasets.","url_abs":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9835116","url_pdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9835116","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"AE-Net (8+16frames)","rank_in_archive_order":29,"of":74,"metrics":{"Top 1 Accuracy":"55.0"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}