{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/glimpse-clouds-human-activity-recognition","title":"Glimpse Clouds: Human Activity Recognition from Unstructured Feature Points","arxiv_id":"1802.07898","date":"2018-02-22","proceeding":"CVPR 2018 6","authors":["Fabien Baradel","Christian Wolf","Julien Mille","Graham W. Taylor"],"abstract":"We propose a method for human activity recognition from RGB data that does\nnot rely on any pose information during test time and does not explicitly\ncalculate pose information internally. Instead, a visual attention module\nlearns to predict glimpse sequences in each frame. These glimpses correspond to\ninterest points in the scene that are relevant to the classified activities. No\nspatial coherence is forced on the glimpse locations, which gives the module\nliberty to explore different points at each frame and better optimize the\nprocess of scrutinizing visual information. Tracking and sequentially\nintegrating this kind of unstructured data is a challenge, which we address by\nseparating the set of glimpses from a set of recurrent tracking/recognition\nworkers. These workers receive glimpses, jointly performing subsequent motion\ntracking and activity prediction. The glimpses are soft-assigned to the\nworkers, optimizing coherence of the assignments in space, time and feature\nspace using an external memory module. No hard decisions are taken, i.e. each\nglimpse point is assigned to all existing workers, albeit with different\nimportance. Our methods outperform state-of-the-art methods on the largest\nhuman activity recognition dataset available to-date; NTU RGB+D Dataset, and on\na smaller human action recognition dataset Northwestern-UCLA Multiview Action\n3D Dataset. Our code is publicly available at\nhttps://github.com/fabienbaradel/glimpse_clouds.","url_abs":"http://arxiv.org/abs/1802.07898v4","url_pdf":"http://arxiv.org/pdf/1802.07898v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"glimpse-clouds-human-activity-recognition","repo_url":"https://github.com/fabienbaradel/glimpse_clouds","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"activity-prediction","task_name":"Activity Prediction"},{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"human-activity-recognition","task_name":"Human Activity Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ntu-rgbd","task":"Action Recognition","dataset":"NTU RGB+D","model":"Glimpse Clouds (RGB only)","rank_in_archive_order":26,"of":28,"metrics":{"Accuracy (CS)":"86.6","Accuracy (CV)":"93.2"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-n-ucla","task":"Skeleton Based Action Recognition","dataset":"N-UCLA","model":"Glimpse Clouds","rank_in_archive_order":25,"of":25,"metrics":{"Accuracy":"87.6%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.07898","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}