{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/finding-action-tubes-with-a-sparse-to-dense","title":"Finding Action Tubes with a Sparse-to-Dense Framework","arxiv_id":"2008.13196","date":"2020-08-30","proceeding":null,"authors":["Yuxi Li","Weiyao Lin","Tao Wang","John See","Rui Qian","Ning Xu","Li-Min Wang","Shugong Xu"],"abstract":"The task of spatial-temporal action detection has attracted increasing attention among researchers. Existing dominant methods solve this problem by relying on short-term information and dense serial-wise detection on each individual frames or clips. Despite their effectiveness, these methods showed inadequate use of long-term information and are prone to inefficiency. In this paper, we propose for the first time, an efficient framework that generates action tube proposals from video streams with a single forward pass in a sparse-to-dense manner. There are two key characteristics in this framework: (1) Both long-term and short-term sampled information are explicitly utilized in our spatiotemporal network, (2) A new dynamic feature sampling module (DTS) is designed to effectively approximate the tube output while keeping the system tractable. We evaluate the efficacy of our model on the UCF101-24, JHMDB-21 and UCFSports benchmark datasets, achieving promising results that are competitive to state-of-the-art methods. The proposed sparse-to-dense strategy rendered our framework about 7.6 times more efficient than the nearest competitor.","url_abs":"https://arxiv.org/abs/2008.13196v1","url_pdf":"https://arxiv.org/pdf/2008.13196v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-detection-on-j-hmdb","task":"Action Detection","dataset":"J-HMDB","model":"DTS","rank_in_archive_order":15,"of":18,"metrics":{"Video-mAP 0.2":"76.1","Video-mAP 0.5":"74.3"},"uses_additional_data":false},{"leaderboard":"/sota/action-detection-on-ucf-sports","task":"Action Detection","dataset":"UCF Sports","model":"DTS","rank_in_archive_order":5,"of":7,"metrics":{"Video-mAP 0.2":"94.3","Video-mAP 0.5":"93.8"},"uses_additional_data":false},{"leaderboard":"/sota/action-detection-on-ucf101-24","task":"Action Detection","dataset":"UCF101-24","model":"DTS","rank_in_archive_order":19,"of":19,"metrics":{"Video-mAP 0.5":"54"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2008.13196","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}