{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/e-2tad-an-energy-efficient-tracking-based","title":"E^2TAD: An Energy-Efficient Tracking-based Action Detector","arxiv_id":"2204.04416","date":"2022-04-09","proceeding":null,"authors":["Xin Hu","Zhenyu Wu","Hao-Yu Miao","Siqi Fan","Taiyu Long","Zhenyu Hu","Pengcheng Pi","Yi Wu","Zhou Ren","Zhangyang Wang","Gang Hua"],"abstract":"Video action detection (spatio-temporal action localization) is usually the starting point for human-centric intelligent analysis of videos nowadays. It has high practical impacts for many applications across robotics, security, healthcare, etc. The two-stage paradigm of Faster R-CNN inspires a standard paradigm of video action detection in object detection, i.e., firstly generating person proposals and then classifying their actions. However, none of the existing solutions could provide fine-grained action detection to the \"who-when-where-what\" level. This paper presents a tracking-based solution to accurately and efficiently localize predefined key actions spatially (by predicting the associated target IDs and locations) and temporally (by predicting the time in exact frame indices). This solution won first place in the UAV-Video Track of 2021 Low-Power Computer Vision Challenge (LPCVC).","url_abs":"https://arxiv.org/abs/2204.04416v4","url_pdf":"https://arxiv.org/pdf/2204.04416v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"e-2tad-an-energy-efficient-tracking-based","repo_url":"https://github.com/VITA-Group/21LPCV-UAV-Solution","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"fine-grained-action-detection","task_name":"Fine-Grained Action Detection"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"spatio-temporal-action-localization","task_name":"Spatio-Temporal Action Localization"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-action-detection","task_name":"Video Action Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"faster-r-cnn","method_name":"Faster R-CNN"},{"method_slug":"rpn","method_name":"RPN"},{"method_slug":"roipool","method_name":"RoIPool"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2204.04416","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}