{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rethinking-the-faster-r-cnn-architecture-for","title":"Rethinking the Faster R-CNN Architecture for Temporal Action Localization","arxiv_id":"1804.07667","date":"2018-04-20","proceeding":"CVPR 2018 6","authors":["Yu-Wei Chao","Sudheendra Vijayanarasimhan","Bryan Seybold","David A. Ross","Jia Deng","Rahul Sukthankar"],"abstract":"We propose TAL-Net, an improved approach to temporal action localization in\nvideo that is inspired by the Faster R-CNN object detection framework. TAL-Net\naddresses three key shortcomings of existing approaches: (1) we improve\nreceptive field alignment using a multi-scale architecture that can accommodate\nextreme variation in action durations; (2) we better exploit the temporal\ncontext of actions for both proposal generation and action classification by\nappropriately extending receptive fields; and (3) we explicitly consider\nmulti-stream feature fusion and demonstrate that fusing motion late is\nimportant. We achieve state-of-the-art performance for both action proposal and\nlocalization on THUMOS'14 detection benchmark and competitive performance on\nActivityNet challenge.","url_abs":"http://arxiv.org/abs/1804.07667v1","url_pdf":"http://arxiv.org/pdf/1804.07667v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"faster-r-cnn","method_name":"Faster R-CNN"},{"method_slug":"rpn","method_name":"RPN"},{"method_slug":"roipool","method_name":"RoIPool"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/temporal-action-localization-on-thumos14","task":"Temporal Action Localization","dataset":"THUMOS’14","model":"TAL-Net","rank_in_archive_order":29,"of":42,"metrics":{"Avg mAP (0.3:0.7)":"39.8","mAP IOU@0.1":"59.8","mAP IOU@0.2":"57.1","mAP IOU@0.3":"53.2","mAP IOU@0.4":"48.5","mAP IOU@0.5":" 42.8","mAP IOU@0.6":"33.8","mAP IOU@0.7":"20.8"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.07667","atlas_url":"https://app.syntology.ai/?focus=1804.07667","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}