Datasets › TED Gesture Dataset
TED Gesture Dataset
Co-speech gestures are everywhere. People make gestures when they chat with others, give a public speech, talk on a phone, and even think aloud. Despite this ubiquity, there are not many datasets available. The main reason is that it is expensive to recruit actors/actresses and track precise body motions. There are a few datasets available (e.g., MSP AVATAR [17] and Personality Dyads Corpus [18]), but their sizes are limited to less than 3 h, and they lack diversity in speech content and speakers. The gestures also could be unnatural owing to inconvenient body tracking suits and acting in a lab environment.
Thus, we collected a new dataset of co-speech gestures: the TED Gesture Dataset. TED is a conference where people share their ideas from a stage, and recordings of these talks are available online. Using TED talks has the following advantages compared to the existing datasets:
• Large enough to learn the mapping from speech to gestures. The number of videos continues to grow. • Various speech content and speakers. There are thousands of unique speakers, and they talk about their own ideas and stories. • The speeches are well prepared, so we expect that the speakers use proper hand gestures. • Favorable for automation of data collection and annotation. All talks come with transcripts, and flat background and steady shots make extracting human poses with computer vision technology easier.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Gesture Generation | TED Gesture Dataset | AQ-GT FGD 1.612 | AQ-GT: a Temporally Aligned and Quantized... | hvoss-techfak/AQGT | 6 | Compare |
Papers archive 2025-07-28
6 shown of 6 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 12. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| AQ-GT: a Temporally Aligned and Quantized GRU-Transformer for Co-Speech Gesture Synthesis | 1 | 1 | 2 May 2023 | not harvested |
| Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings | 1 | 1 | 4 Oct 2022 | not harvested |
| Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation | 1 | 1 | 24 Mar 2022 | not harvested |
| SEEG: Semantic Energized Co-Speech Gesture Generation | 1 | 1 | 1 Jan 2022 | not harvested |
| Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning | 1 | 1 | 31 Jul 2021 | not harvested |
| Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity | 2 | 1 | 4 Sep 2020 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- TED Gesture Dataset
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections