{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/review-of-action-recognition-and-detection","title":"Review of Action Recognition and Detection Methods","arxiv_id":"1610.06906","date":"2016-10-21","proceeding":null,"authors":["Soo Min Kang","Richard P. Wildes"],"abstract":"In computer vision, action recognition refers to the act of classifying an\naction that is present in a given video and action detection involves locating\nactions of interest in space and/or time. Videos, which contain photometric\ninformation (e.g. RGB, intensity values) in a lattice structure, contain\ninformation that can assist in identifying the action that has been imaged. The\nprocess of action recognition and detection often begins with extracting useful\nfeatures and encoding them to ensure that the features are specific to serve\nthe task of action recognition and detection. Encoded features are then\nprocessed through a classifier to identify the action class and their spatial\nand/or temporal locations. In this report, a thorough review of various action\nrecognition and detection algorithms in computer vision is provided by\nanalyzing the two-step process of a typical action recognition and detection\nalgorithm: (i) extraction and encoding of features, and (ii) classifying\nfeatures into action classes. In efforts to ensure that computer vision-based\nalgorithms reach the capabilities that humans have of identifying actions\nirrespective of various nuisance variables that may be present within the field\nof view, the state-of-the-art methods are reviewed and some remaining problems\nare addressed in the final chapter.","url_abs":"http://arxiv.org/abs/1610.06906v2","url_pdf":"http://arxiv.org/pdf/1610.06906v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"review-of-action-recognition-and-detection","repo_url":"https://github.com/sulasen/race-events-recognition-1","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1610.06906","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}