Browse State-of-the-Art › Gesture Generation
Gesture Generation
46 papers with code · 4 benchmarks · 7 datasets archive 2025-07-28
Generation of gestures, as a sequence of 3d poses
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| BEAT2 (14 rows) | Intentional Gesture | Intentional Gesture: Deliver Your Intentions with Gestures for Speech | code | — | Compare |
| TED Gesture Dataset (6 rows) | AQ-GT | AQ-GT: a Temporally Aligned and Quantized GRU-Transformer for... | code | — | Compare |
| BEAT (5 rows) | CaMN | BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for... | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| DVS128 Gesture (1 row) | ECSNet | Ecsnet: Spatio-temporal feature learning for event camera | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 46 papers with code (107 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Sep 2020 7 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 1 pointer-only (licence)robosuite is a simulation framework for robot learning powered by the MuJoCo physics engine.
-
8 Dec 2022 3 repositories listedThis work addresses the problem of generating 3D holistic body motions from human speech.
-
22 Aug 2022 3 repositories listedOn the other hand, all synthetic motion is found to be vastly less appropriate for the speech than the original motion-capture recordings.
-
24 Aug 2023 2 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedThe effect of the interlocutor is even more subtle, with submitted systems at best performing barely above chance.
-
4 Sep 2020 2 repositories listedIn this paper, we present an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures.
-
10 Jun 2019 2 repositories listedSpecifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a single speaker to their hand and arm motion.
-
3 Jul 2025 1 repository listedAlong with the explosion of large language models, improvements in speech synthesis, advancements in hardware, and the evolution of computer graphics, the current bottleneck in creating digital humans lies in generating…
-
21 May 2025 1 repository listedTo address this gap, we introduce \textbf{Intentional-Gesture}, a novel framework that casts gesture generation as an intention-reasoning task grounded in high-level communicative functions.
-
15 Feb 2025 1 repository listedGenerating expressive and diverse human gestures from audio is crucial in fields like human-computer interaction, virtual reality, and animation.
-
31 Jan 2025 1 repository listedTo overcome the suboptimal performance of flow matching baseline, we propose latent shortcut learning and beta distribution time stamp sampling during training to enhance gesture synthesis quality and accelerate…
-
9 Dec 2024 1 repository listedTherefore, we present RAG-Gesture, a diffusion-based gesture generation approach that leverages Retrieval Augmented Generation (RAG) to produce natural-looking and semantically rich gestures.
-
1 Oct 2024 1 repository listedCurrent co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts,…
-
30 Jul 2024 1 repository listedHowever, employing a unified model to achieve various generation tasks with different condition modalities presents two main challenges: motion distribution drifts across different tasks (e.
-
1 Jun 2024 1 repository listedOnce trained, AMUSE synthesizes 3D human gestures directly from speech with control over the expressed emotions and style by combining the content from the driving speech with the emotion and style of another speech…
-
26 Mar 2024 1 repository listedGestures play a key role in human communication.
-
14 Mar 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedGesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality.
-
31 Dec 2023 1 repository listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements.
-
7 Dec 2023 1 repository listedOnce trained, AMUSE synthesizes 3D human gestures directly from speech with control over the expressed emotions and style by combining the content from the driving speech with the emotion and style of another speech…
-
17 Sep 2023 1 repository listed Syntology ran 8 of 10 samples · 2 unverified · 10 pointer-only (licence)While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in the gesticulations.
-
13 Sep 2023 1 repository listedThe automatic co-speech gesture generation draws much attention in computer animation.
-
8 Sep 2023 1 repository listedCreating a diverse and comprehensive dataset of hand gestures for dynamic human-machine interfaces in the automotive domain can be challenging and time-consuming.
-
29 Aug 2023 1 repository listedCo-speech gesture generation is crucial for automatic digital avatar animation.
-
30 May 2023 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)In this work, we propose EmotionGesture, a novel framework for synthesizing vivid and diverse emotional co-speech 3D gestures from audio.
-
18 May 2023 1 repository listed Syntology ran 3 of 5 samples · 2 unverified · 5 pointer-only (licence)Levenshtein distance based on audio quantization as a similarity metric of corresponding speech of gestures helps match more appropriate gestures with speech, and solves the alignment problem of speech and gestures well.
-
2 May 2023 1 repository listedBy learning the mapping of a latent space representation as opposed to directly mapping it to a vector representation, this framework facilitates the generation of highly realistic and expressive gestures that closely…
-
26 Mar 2023 1 repository listedWe leverage the power of the large-scale Contrastive-Language-Image-Pre-training (CLIP) model and present a novel CLIP-guided mechanism that extracts efficient style representations from multiple input modalities, such…
-
16 Mar 2023 1 repository listedIn this work, we propose a novel diffusion-based framework, named Diffusion Co-Speech Gesture (DiffGesture), to effectively capture the cross-modal audio-to-gesture associations and preserve temporal coherence for…
-
17 Nov 2022 1 repository listedDiffusion models have experienced a surge of interest as highly expressive yet efficiently trainable probabilistic models.
-
4 Oct 2022 1 repository listedWe present a novel co-speech gesture synthesis method that achieves convincing results both on the rhythm and semantics.
-
15 Sep 2022 1 repository listedIn a series of experiments, we first demonstrate the flexibility and generalizability of our model to new speakers and styles.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections