{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-feature-learning-of-human","title":"Unsupervised Feature Learning of Human Actions as Trajectories in Pose Embedding Manifold","arxiv_id":"1812.02592","date":"2018-12-06","proceeding":null,"authors":["Jogendra Nath Kundu","Maharshi Gor","Phani Krishna Uppala","R. Venkatesh Babu"],"abstract":"An unsupervised human action modeling framework can provide useful\npose-sequence representation, which can be utilized in a variety of pose\nanalysis applications. In this work we propose a novel temporal pose-sequence\nmodeling framework, which can embed the dynamics of 3D human-skeleton joints to\na continuous latent space in an efficient manner. In contrast to end-to-end\nframework explored by previous works, we disentangle the task of individual\npose representation learning from the task of learning actions as a trajectory\nin pose embedding space. In order to realize a continuous pose embedding\nmanifold with improved reconstructions, we propose an unsupervised, manifold\nlearning procedure named Encoder GAN, (or EnGAN). Further, we use the pose\nembeddings generated by EnGAN to model human actions using a bidirectional RNN\nauto-encoder architecture, PoseRNN. We introduce first-order gradient loss to\nexplicitly enforce temporal regularity in the predicted motion sequence. A\nhierarchical feature fusion technique is also investigated for simultaneous\nmodeling of local skeleton joints along with global pose variations. We\ndemonstrate state-of-the-art transfer-ability of the learned representation\nagainst other supervisedly and unsupervisedly learned motion embeddings for the\ntask of fine-grained action recognition on SBU interaction dataset. Further, we\nshow the qualitative strengths of the proposed framework by visualizing\nskeleton pose reconstructions and interpolations in pose-embedding space, and\nlow dimensional principal component projections of the reconstructed pose\ntrajectories.","url_abs":"http://arxiv.org/abs/1812.02592v1","url_pdf":"http://arxiv.org/pdf/1812.02592v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-feature-learning-of-human","repo_url":"https://github.com/ThomasDupiereux/ADLproject","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"unsupervised-feature-learning-of-human","repo_url":"https://github.com/ThomasDupiereux/EnGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"unsupervised-feature-learning-of-human","repo_url":"https://github.com/maharshi95/Pose2vec","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"fine-grained-action-recognition","task_name":"Fine-grained Action Recognition"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1812.02592","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}