{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lets-dance-learning-from-online-dance-videos","title":"Let's Dance: Learning From Online Dance Videos","arxiv_id":"1801.07388","date":"2018-01-23","proceeding":null,"authors":["Daniel Castro","Steven Hickson","Patsorn Sangkloy","Bhavishya Mittal","Sean Dai","James Hays","Irfan Essa"],"abstract":"In recent years, deep neural network approaches have naturally extended to\nthe video domain, in their simplest case by aggregating per-frame\nclassifications as a baseline for action recognition. A majority of the work in\nthis area extends from the imaging domain, leading to visual-feature heavy\napproaches on temporal data. To address this issue we introduce \"Let's Dance\",\na 1000 video dataset (and growing) comprised of 10 visually overlapping dance\ncategories that require motion for their classification. We stress the\nimportant of human motion as a key distinguisher in our work given that, as we\nshow in this work, visual information is not sufficient to classify\nmotion-heavy categories. We compare our datasets' performance using imaging\ntechniques with UCF-101 and demonstrate this inherent difficulty. We present a\ncomparison of numerous state-of-the-art techniques on our dataset using three\ndifferent representations (video, optical flow and multi-person pose data) in\norder to analyze these approaches. We discuss the motion parameterization of\neach of them and their value in learning to categorize online dance videos.\nLastly, we release this dataset (and its three representations) for the\nresearch community to use.","url_abs":"http://arxiv.org/abs/1801.07388v1","url_pdf":"http://arxiv.org/pdf/1801.07388v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lets-dance-learning-from-online-dance-videos","repo_url":"https://github.com/xrenaa/Human-Motion-Analysis-with-Deep-Metric-Learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1801.07388","atlas_url":"https://app.syntology.ai/?focus=1801.07388","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}