{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/real-time-action-detection-in-video","title":"Real-Time Action Detection in Video Surveillance using Sub-Action Descriptor with Multi-CNN","arxiv_id":"1710.03383","date":"2017-10-10","proceeding":null,"authors":["Cheng-Bin Jin","Shengzhe Li","Hakil Kim"],"abstract":"When we say a person is texting, can you tell the person is walking or\nsitting? Emphatically, no. In order to solve this incomplete representation\nproblem, this paper presents a sub-action descriptor for detailed action\ndetection. The sub-action descriptor consists of three levels: the posture, the\nlocomotion, and the gesture level. The three levels give three sub-action\ncategories for one action to address the representation problem. The proposed\naction detection model simultaneously localizes and recognizes the actions of\nmultiple individuals in video surveillance using appearance-based temporal\nfeatures with multi-CNN. The proposed approach achieved a mean average\nprecision (mAP) of 76.6% at the frame-based and 83.5% at the video-based\nmeasurement on the new large-scale ICVL video surveillance dataset that the\nauthors introduce and make available to the community with this paper.\nExtensive experiments on the benchmark KTH dataset demonstrate that the\nproposed approach achieved better performance, which in turn boosts the action\nrecognition performance over the state-of-the-art. The action detection model\ncan run at around 25 fps on the ICVL and more than 80 fps on the KTH dataset,\nwhich is suitable for real-time surveillance applications.","url_abs":"http://arxiv.org/abs/1710.03383v1","url_pdf":"http://arxiv.org/pdf/1710.03383v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"real-time-action-detection-in-video","repo_url":"https://github.com/ChengBinJin/ActionViewer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1710.03383","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}