Datasets › BEAT
BEAT (Body-Expression-Audio-Text)
BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations. Our statistical analysis on BEAT demonstrates the correlation of conversational gestures with \textit{facial expressions}, \textit{emotions}, and \textit{semantics}, in addition to the known correlation with \textit{audio}, \textit{text}, and \textit{speaker identity}. Based on this observation, we propose a baseline model, \textbf{Ca}scaded \textbf{M}otion \textbf{N}etwork \textbf{(CaMN)}, which consists of above six modalities modeled in a cascaded architecture for gesture synthesis. To evaluate the semantic relevancy, we introduce a metric, Semantic Relevance Gesture Recall (\textbf{SRGR}). Qualitative and quantitative experiments demonstrate metrics' validness, ground truth data quality, and baseline's state-of-the-art performance. To the best of our knowledge, BEAT is the largest motion capture dataset for investigating human gestures, which may contribute to a number of different research fields, including controllable gesture synthesis, cross-modality analysis, and emotional gesture recognition. The data, code and model are available on \url{https://pantomatrix.github.io/BEAT/}.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Gesture Generation | BEAT | CaMN FID 122.8 | BEAT: A Large-Scale Semantic and Emotional Multi-Modal... | PantoMatrix/PantoMatrix | 5 | Compare |
Papers archive 2025-07-28
5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 56. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis | 1 | 1 | 10 Mar 2022 | ran 3 of 3 samples (0 unverified; 3 pointer-only for licence) |
| Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders | 0 | 1 | 15 Aug 2021 | not harvested |
| Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity | 2 | 1 | 4 Sep 2020 | not harvested |
| Learning Individual Styles of Conversational Gesture | 2 | 1 | 10 Jun 2019 | not harvested |
| Robots Learning to Say `No': Prohibition and Rejective Mechanisms in Acquisition of Linguistic Negation | 0 | 1 | 28 Oct 2018 | not harvested |
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- BEAT
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections