{"url":"/dataset/aist","name":"AIST++","full_name":null,"description_markdown":"**AIST++** is a 3D dance dataset which contains 3D motion reconstructed from real dancers paired with music. The AIST++ Dance Motion Dataset is constructed from the AIST Dance Video DB. With multi-view videos, an elaborate pipeline is designed to estimate the camera parameters, 3D human keypoints and 3D human dance motion sequences:\r\n\r\n- It provides 3D human keypoint annotations and camera parameters for 10.1M images, covering 30 different subjects in 9 views. These attributes makes it the largest and richest existing dataset with 3D human keypoint annotations.\r\n- It also contains 1,408 sequences of 3D human dance motion, represented as joint rotations along with root trajectories. The dance motions are equally distributed among 10 dance genres with hundreds of choreographies. Motion durations vary from 7.4 sec. to 48.0 sec. All the dance motions have corresponding music.","description_withheld":null,"homepage":"https://google.github.io/aistplusplus_dataset/","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"3D","url":"/datasets/modality/3d"}],"tasks":[{"name":"3D Human Pose Estimation","url":"/task/3d-human-pose-estimation","datasets_with_task":"/datasets/task/3d-human-pose-estimation"},{"name":"Motion Synthesis","url":"/task/motion-synthesis","datasets_with_task":"/datasets/task/motion-synthesis"}],"languages":[],"variants":["AIST++"],"data_loaders":[],"num_papers_in_archive":21,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/motion-synthesis-on-aist","task":"Motion Synthesis","dataset_variant":"AIST++","rows":12,"metrics":["FID","Beat alignment score"],"first_row_in_archive_order":{"model":"Motion Anything","paper":"/paper/motion-anything-any-to-motion-generation","metrics":{"Beat alignment score":"0.2757","FID":"17.22"},"code_links":[{"title":"steve-zeyu-zhang/MotionAnything","url":"https://github.com/steve-zeyu-zhang/MotionAnything"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/3d-human-pose-estimation-on-aist","task":"3D Human Pose Estimation","dataset_variant":"AIST++","rows":5,"metrics":["MPJPE","Single-view","Acceleration Error"],"first_row_in_archive_order":{"model":"RobustCap","paper":"/paper/fusing-monocular-images-and-sparse-imu","metrics":{"MPJPE":"33.1"},"code_links":[{"title":"shaohua-pan/RobustCap","url":"https://github.com/shaohua-pan/RobustCap"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/motion-anything-any-to-motion-generation","title":"Motion Anything: Any to Motion Generation","date":"2025-03-10","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/lodge-a-coarse-to-fine-diffusion-network-for","title":"Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives","date":"2024-03-15","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":17,"samples_ran":10,"samples_unverified":7,"pointer_only_for_licence":17,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/fusing-monocular-images-and-sparse-imu","title":"Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion Capture","date":"2023-09-01","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/tm2d-bimodality-driven-3d-dance-generation","title":"TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration","date":"2023-04-05","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mofusion-a-framework-for-denoising-diffusion","title":"MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis","date":"2022-12-08","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/ude-a-unified-driving-engine-for-human-motion","title":"UDE: A Unified Driving Engine for Human Motion Generation","date":"2022-11-29","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/kinematic-aware-hierarchical-attention","title":"Kinematic-aware Hierarchical Attention Network for Human Pose Estimation in Videos","date":"2022-11-29","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/edge-editable-dance-generation-from-music","title":"EDGE: Editable Dance Generation From Music","date":"2022-11-19","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":11,"samples_ran":8,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/bailando-3d-dance-generation-by-actor-critic","title":"Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory","date":"2022-03-24","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":5,"samples_ran":4,"samples_unverified":1,"pointer_only_for_licence":5,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/hybridcap-inertia-aid-monocular-capture-of","title":"HybridCap: Inertia-aid Monocular Capture of Challenging Human Motions","date":"2022-03-17","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/deciwatch-a-simple-baseline-for-10x-efficient","title":"DeciWatch: A Simple Baseline for 10x Efficient 2D and 3D Pose Estimation","date":"2022-03-16","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":8,"samples_ran":6,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learn-to-dance-with-aist-music-conditioned-3d","title":"AI Choreographer: Music Conditioned 3D Dance Generation with AIST++","date":"2021-01-21","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/dance-revolution-long-sequence-dance","title":"Dance Revolution: Long-Term Dance Generation with Music via Curriculum Learning","date":"2020-06-11","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/music2dance-music-driven-dance-generation","title":"Music2Dance: DanceNet for Music-driven Dance Generation","date":"2020-02-02","rows_on_this_dataset":1,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":5,"samples_harvested":42,"samples_ran":29,"samples_unverified":13,"pointer_only_for_licence":23,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}