{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recurrent-network-models-for-human-dynamics","title":"Recurrent Network Models for Human Dynamics","arxiv_id":"1508.00271","date":"2015-08-02","proceeding":"ICCV 2015 12","authors":["Katerina Fragkiadaki","Sergey Levine","Panna Felsen","Jitendra Malik"],"abstract":"We propose the Encoder-Recurrent-Decoder (ERD) model for recognition and\nprediction of human body pose in videos and motion capture. The ERD model is a\nrecurrent neural network that incorporates nonlinear encoder and decoder\nnetworks before and after recurrent layers. We test instantiations of ERD\narchitectures in the tasks of motion capture (mocap) generation, body pose\nlabeling and body pose forecasting in videos. Our model handles mocap training\ndata across multiple subjects and activity domains, and synthesizes novel\nmotions while avoid drifting for long periods of time. For human pose labeling,\nERD outperforms a per frame body part detector by resolving left-right body\npart confusions. For video pose forecasting, ERD predicts body joint\ndisplacements across a temporal horizon of 400ms and outperforms a first order\nmotion model based on optical flow. ERDs extend previous Long Short Term Memory\n(LSTM) models in the literature to jointly learn representations and their\ndynamics. Our experiments show such representation learning is crucial for both\nlabeling and prediction in space-time. We find this is a distinguishing feature\nbetween the spatio-temporal visual domain in comparison to 1D text, speech or\nhandwriting, where straightforward hard coded representations have shown\nexcellent results when directly combined with recurrent units.","url_abs":"http://arxiv.org/abs/1508.00271v2","url_pdf":"http://arxiv.org/pdf/1508.00271v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"human-dynamics","task_name":"Human Dynamics"},{"task_slug":"human-pose-forecasting","task_name":"Human Pose Forecasting"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-pose-forecasting-on-human36m","task":"Human Pose Forecasting","dataset":"Human3.6M","model":"ERD","rank_in_archive_order":21,"of":33,"metrics":{"MAR, walking, 1,000ms":"2.38","MAR, walking, 400ms":"1.78"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}