{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/decoupling-localization-and-classification-in","title":"Decoupling Localization and Classification in Single Shot Temporal Action Detection","arxiv_id":"1904.07442","date":"2019-04-16","proceeding":null,"authors":["Yupan Huang","Qi Dai","Yutong Lu"],"abstract":"Video temporal action detection aims to temporally localize and recognize the\naction in untrimmed videos. Existing one-stage approaches mostly focus on\nunifying two subtasks, i.e., localization of action proposals and\nclassification of each proposal through a fully shared backbone. However, such\ndesign of encapsulating all components of two subtasks in one single network\nmight restrict the training by ignoring the specialized characteristic of each\nsubtask. In this paper, we propose a novel Decoupled Single Shot temporal\nAction Detection (Decouple-SSAD) method to mitigate such problem by decoupling\nthe localization and classification in a one-stage scheme. Particularly, two\nseparate branches are designed in parallel to enable each component to own\nrepresentations privately for accurate localization or classification. Each\nbranch produces a set of action anchor layers by applying deconvolution to the\nfeature maps of the main stream. Each branch produces a set of feature maps by\napplying deconvolution to the feature maps of the main stream. High-level\nsemantic information from deeper layers is thus incorporated to enhance the\nfeature representations. We conduct extensive experiments on THUMOS14 dataset\nand demonstrate superior performance over state-of-the-art methods. Our code is\navailable online.","url_abs":"http://arxiv.org/abs/1904.07442v1","url_pdf":"http://arxiv.org/pdf/1904.07442v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"decoupling-localization-and-classification-in","repo_url":"https://github.com/hypjudy/Decouple-SSAD","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/temporal-action-localization-on-thumos14","task":"Temporal Action Localization","dataset":"THUMOS’14","model":"Decouple-SSAD","rank_in_archive_order":28,"of":42,"metrics":{"Avg mAP (0.3:0.7)":"42.0","mAP IOU@0.3":"60.2","mAP IOU@0.4":"54.1","mAP IOU@0.5":"44.2","mAP IOU@0.6":"32.3","mAP IOU@0.7":"19.1"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1904.07442","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}