{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/star-net-action-recognition-using-spatio","title":"STAR-Net: Action Recognition using Spatio-Temporal Activation Reprojection","arxiv_id":"1902.10024","date":"2019-02-26","proceeding":null,"authors":["William McNally","Alexander Wong","John McPhee"],"abstract":"While depth cameras and inertial sensors have been frequently leveraged for\nhuman action recognition, these sensing modalities are impractical in many\nscenarios where cost or environmental constraints prohibit their use. As such,\nthere has been recent interest on human action recognition using low-cost,\nreadily-available RGB cameras via deep convolutional neural networks. However,\nmany of the deep convolutional neural networks proposed for action recognition\nthus far have relied heavily on learning global appearance cues directly from\nimaging data, resulting in highly complex network architectures that are\ncomputationally expensive and difficult to train. Motivated to reduce network\ncomplexity and achieve higher performance, we introduce the concept of\nspatio-temporal activation reprojection (STAR). More specifically, we reproject\nthe spatio-temporal activations generated by human pose estimation layers in\nspace and time using a stack of 3D convolutions. Experimental results on\nUTD-MHAD and J-HMDB demonstrate that an end-to-end architecture based on the\nproposed STAR framework (which we nickname STAR-Net) is proficient in\nsingle-environment and small-scale applications. On UTD-MHAD, STAR-Net\noutperforms several methods using richer data modalities such as depth and\ninertial sensors.","url_abs":"http://arxiv.org/abs/1902.10024v1","url_pdf":"http://arxiv.org/pdf/1902.10024v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"multimodal-activity-recognition","task_name":"Multimodal Activity Recognition"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multimodal-activity-recognition-on-utd-mhad","task":"Multimodal Activity Recognition","dataset":"UTD-MHAD","model":"STAR-Net","rank_in_archive_order":5,"of":5,"metrics":{"Accuracy (CS)":"90"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-j-hmdb","task":"Skeleton Based Action Recognition","dataset":"J-HMDB","model":"STAR-Net","rank_in_archive_order":9,"of":13,"metrics":{"Accuracy (RGB+pose)":"64.3"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}