{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/finding-action-tubes","title":"Finding Action Tubes","arxiv_id":"1411.6031","date":"2014-11-21","proceeding":"CVPR 2015 6","authors":["Georgia Gkioxari","Jitendra Malik"],"abstract":"We address the problem of action detection in videos. Driven by the latest\nprogress in object detection from 2D images, we build action models using rich\nfeature hierarchies derived from shape and kinematic cues. We incorporate\nappearance and motion in two ways. First, starting from image region proposals\nwe select those that are motion salient and thus are more likely to contain the\naction. This leads to a significant reduction in the number of regions being\nprocessed and allows for faster computations. Second, we extract\nspatio-temporal feature representations to build strong classifiers using\nConvolutional Neural Networks. We link our predictions to produce detections\nconsistent in time, which we call action tubes. We show that our approach\noutperforms other techniques in the task of action detection.","url_abs":"http://arxiv.org/abs/1411.6031v1","url_pdf":"http://arxiv.org/pdf/1411.6031v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"finding-action-tubes","repo_url":"https://github.com/JeffCHEN2017/WSSTG","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-detection-on-j-hmdb","task":"Action Detection","dataset":"J-HMDB","model":"Action Tubes","rank_in_archive_order":13,"of":18,"metrics":{"Frame-mAP 0.5":"36.2","Video-mAP 0.5":"53.3"},"uses_additional_data":false},{"leaderboard":"/sota/action-detection-on-ucf-sports","task":"Action Detection","dataset":"UCF Sports","model":"Action Tubes","rank_in_archive_order":4,"of":7,"metrics":{"Frame-mAP 0.5":"68.1","Video-mAP 0.5":"75.8"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-j-hmdb","task":"Skeleton Based Action Recognition","dataset":"J-HMDB","model":"Action Tubes","rank_in_archive_order":10,"of":13,"metrics":{"Accuracy (RGB+pose)":"62.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1411.6031","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}