{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/online-human-action-detection-using-joint","title":"Online Human Action Detection using Joint Classification-Regression Recurrent Neural Networks","arxiv_id":"1604.05633","date":"2016-04-19","proceeding":null,"authors":["Yanghao Li","Cuiling Lan","Junliang Xing","Wen-Jun Zeng","Chunfeng Yuan","Jiaying Liu"],"abstract":"Human action recognition from well-segmented 3D skeleton data has been\nintensively studied and has been attracting an increasing attention. Online\naction detection goes one step further and is more challenging, which\nidentifies the action type and localizes the action positions on the fly from\nthe untrimmed stream data. In this paper, we study the problem of online action\ndetection from streaming skeleton data. We propose a multi-task end-to-end\nJoint Classification-Regression Recurrent Neural Network to better explore the\naction type and temporal localization information. By employing a joint\nclassification and regression optimization objective, this network is capable\nof automatically localizing the start and end points of actions more\naccurately. Specifically, by leveraging the merits of the deep Long Short-Term\nMemory (LSTM) subnetwork, the proposed model automatically captures the complex\nlong-range temporal dynamics, which naturally avoids the typical sliding window\ndesign and thus ensures high computational efficiency. Furthermore, the subtask\nof regression optimization provides the ability to forecast the action prior to\nits occurrence. To evaluate our proposed model, we build a large streaming\nvideo dataset with annotations. Experimental results on our dataset and the\npublic G3D dataset both demonstrate very promising performance of our scheme.","url_abs":"http://arxiv.org/abs/1604.05633v2","url_pdf":"http://arxiv.org/pdf/1604.05633v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"online-human-action-detection-using-joint","repo_url":"https://github.com/seanmcgovern21/Machine-Learning-CS539","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"online-action-detection","task_name":"Online Action Detection"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"temporal-localization","task_name":"Temporal Localization"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1604.05633","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}