{"url":"/dataset/beat","name":"BEAT","full_name":"Body-Expression-Audio-Text","description_markdown":"BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations.\r\nOur statistical analysis on BEAT demonstrates the correlation of conversational gestures with \\textit{facial expressions}, \\textit{emotions}, and \\textit{semantics}, in addition to the known correlation with \\textit{audio}, \\textit{text}, and \\textit{speaker identity}.\r\nBased on this observation, we propose a baseline model, \\textbf{Ca}scaded \\textbf{M}otion \\textbf{N}etwork \\textbf{(CaMN)}, which consists of above six modalities modeled in a cascaded architecture for gesture synthesis. To evaluate the semantic relevancy, we introduce a metric, Semantic Relevance Gesture Recall (\\textbf{SRGR}). \r\nQualitative and quantitative experiments demonstrate metrics' validness, ground truth data quality, and baseline's state-of-the-art performance. \r\nTo the best of our knowledge, BEAT is the largest motion capture dataset for investigating human gestures, which may contribute to a number of different research fields, including controllable gesture synthesis, cross-modality analysis, and emotional gesture recognition. The data, code and model are available on \\url{https://pantomatrix.github.io/BEAT/}.","description_withheld":null,"homepage":"https://pantomatrix.github.io/BEAT/","introduced_date":"2022-03-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/beat-a-large-scale-semantic-and-emotional","title":"BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis","first_author":"Haiyang Liu","url":null},"license":{"name":"CC-BY-NC-4.0","url":"https://github.com/PantoMatrix/BEAT/blob/main/LICENSE.md"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"3D","url":"/datasets/modality/3d"},{"name":"Audio","url":"/datasets/modality/audio"},{"name":"3d meshes","url":"/datasets/modality/3d-meshes"},{"name":"Speech","url":"/datasets/modality/speech"},{"name":"Actions","url":"/datasets/modality/actions"}],"tasks":[{"name":"Gesture Generation","url":"/task/gesture-generation","datasets_with_task":"/datasets/task/gesture-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"Spanish","url":"/datasets/language/spanish"},{"name":"Chinese","url":"/datasets/language/chinese"},{"name":"Japanese","url":"/datasets/language/japanese"}],"variants":["BEAT"],"data_loaders":[{"repo":"https://github.com/PantoMatrix/PantoMatrix","url":"https://github.com/PantoMatrix/PantoMatrix/blob/main/scripts/BEAT_2022/readme_beat.md","frameworks":["pytorch"]}],"num_papers_in_archive":56,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/gesture-generation-on-beat","task":"Gesture Generation","dataset_variant":"BEAT","rows":5,"metrics":["FID"],"first_row_in_archive_order":{"model":"CaMN","paper":"/paper/beat-a-large-scale-semantic-and-emotional","metrics":{"FID":"122.8"},"code_links":[{"title":"PantoMatrix/PantoMatrix","url":"https://github.com/PantoMatrix/PantoMatrix"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/beat-a-large-scale-semantic-and-emotional","title":"BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis","date":"2022-03-10","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/audio2gestures-generating-diverse-gestures","title":"Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders","date":"2021-08-15","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/speech-gesture-generation-from-the-trimodal","title":"Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity","date":"2020-09-04","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/learning-individual-styles-of-conversational-1","title":"Learning Individual Styles of Conversational Gesture","date":"2019-06-10","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/robots-learning-to-say-no-prohibition-and","title":"Robots Learning to Say `No': Prohibition and Rejective Mechanisms in Acquisition of Linguistic Negation","date":"2018-10-28","rows_on_this_dataset":1,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":3,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}