{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-large-scale-varying-view-rgb-d-action","title":"A Large-scale Varying-view RGB-D Action Dataset for Arbitrary-view Human Action Recognition","arxiv_id":"1904.10681","date":"2019-04-24","proceeding":null,"authors":["Yanli Ji","Feixiang Xu","Yang Yang","Fumin Shen","Heng Tao Shen","Wei-Shi Zheng"],"abstract":"Current researches of action recognition mainly focus on single-view and\nmulti-view recognition, which can hardly satisfies the requirements of\nhuman-robot interaction (HRI) applications to recognize actions from arbitrary\nviews. The lack of datasets also sets up barriers. To provide data for\narbitrary-view action recognition, we newly collect a large-scale RGB-D action\ndataset for arbitrary-view action analysis, including RGB videos, depth and\nskeleton sequences. The dataset includes action samples captured in 8 fixed\nviewpoints and varying-view sequences which covers the entire 360 degree view\nangles. In total, 118 persons are invited to act 40 action categories, and\n25,600 video samples are collected. Our dataset involves more participants,\nmore viewpoints and a large number of samples. More importantly, it is the\nfirst dataset containing the entire 360 degree varying-view sequences. The\ndataset provides sufficient data for multi-view, cross-view and arbitrary-view\naction analysis. Besides, we propose a View-guided Skeleton CNN (VS-CNN) to\ntackle the problem of arbitrary-view action recognition. Experiment results\nshow that the VS-CNN achieves superior performance.","url_abs":"http://arxiv.org/abs/1904.10681v1","url_pdf":"http://arxiv.org/pdf/1904.10681v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-analysis","task_name":"Action Analysis"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[{"slug":"uestc-rgb-d","name":"UESTC RGB-D","full_name":"UESTC RGB-D Varying-view action database"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-varying","task":"Skeleton Based Action Recognition","dataset":"Varying-view RGB-D Action-Skeleton","model":"VS-CNN","rank_in_archive_order":1,"of":7,"metrics":{"Accuracy (AV I)":"57%","Accuracy (AV II)":"75%","Accuracy (CS)":"76%","Accuracy (CV I)":"29%","Accuracy (CV II)":"71%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.10681","atlas_url":"https://app.syntology.ai/?focus=1904.10681","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}