{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/from-pose-to-activity-surveying-datasets-and","title":"From Pose to Activity: Surveying Datasets and Introducing CONVERSE","arxiv_id":"1511.05788","date":"2015-11-18","proceeding":null,"authors":["Michael Edwards","Jingjing Deng","Xianghua Xie"],"abstract":"We present a review on the current state of publicly available datasets\nwithin the human action recognition community; highlighting the revival of pose\nbased methods and recent progress of understanding person-person interaction\nmodeling. We categorize datasets regarding several key properties for usage as\na benchmark dataset; including the number of class labels, ground truths\nprovided, and application domain they occupy. We also consider the level of\nabstraction of each dataset; grouping those that present actions, interactions\nand higher level semantic activities. The survey identifies key appearance and\npose based datasets, noting a tendency for simplistic, emphasized, or scripted\naction classes that are often readily definable by a stable collection of\nsub-action gestures. There is a clear lack of datasets that provide closely\nrelated actions, those that are not implicitly identified via a series of poses\nand gestures, but rather a dynamic set of interactions. We therefore propose a\nnovel dataset that represents complex conversational interactions between two\nindividuals via 3D pose. 8 pairwise interactions describing 7 separate\nconversation based scenarios were collected using two Kinect depth sensors. The\nintention is to provide events that are constructed from numerous primitive\nactions, interactions and motions, over a period of time; providing a set of\nsubtle action classes that are more representative of the real world, and a\nchallenge to currently developed recognition methodologies. We believe this is\namong one of the first datasets devoted to conversational interaction\nclassification using 3D pose features and the attributed papers show this task\nis indeed possible. The full dataset is made publicly available to the research\ncommunity at www.csvision.swansea.ac.uk/converse.","url_abs":"http://arxiv.org/abs/1511.05788v2","url_pdf":"http://arxiv.org/pdf/1511.05788v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[{"slug":"converse","name":"CONVERSE","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}