{"url":"/dataset/ted-gesture-dataset","name":"TED Gesture Dataset","full_name":null,"description_markdown":"Co-speech gestures are everywhere. People make gestures\r\nwhen they chat with others, give a public speech, talk on a\r\nphone, and even think aloud. Despite this ubiquity, there are\r\nnot many datasets available. The main reason is that it is\r\nexpensive to recruit actors/actresses and track precise body\r\nmotions. There are a few datasets available (e.g., MSP\r\nAVATAR [17] and Personality Dyads Corpus [18]), but their\r\nsizes are limited to less than 3 h, and they lack diversity in\r\nspeech content and speakers. The gestures also could be\r\nunnatural owing to inconvenient body tracking suits and acting\r\nin a lab environment.\r\n\r\nThus, we collected a new dataset of co-speech gestures: the\r\nTED Gesture Dataset. TED is a conference where people share\r\ntheir ideas from a stage, and recordings of these talks are\r\navailable online. Using TED talks has the following\r\nadvantages compared to the existing datasets:\r\n\r\n• Large enough to learn the mapping from speech to\r\ngestures. The number of videos continues to grow.\r\n• Various speech content and speakers. There are\r\nthousands of unique speakers, and they talk about their\r\nown ideas and stories.\r\n• The speeches are well prepared, so we expect that the\r\nspeakers use proper hand gestures.\r\n• Favorable for automation of data collection and\r\nannotation. All talks come with transcripts, and flat\r\nbackground and steady shots make extracting human\r\nposes with computer vision technology easier.","description_withheld":null,"homepage":"https://github.com/youngwoo-yoon/youtube-gesture-dataset","introduced_date":"2018-10-30","introduced_date_note":null,"introduced_by":{"paper":"/paper/robots-learn-social-skills-end-to-end","title":"Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots","first_author":null,"url":null},"license":{"name":"BSD-3","url":"https://github.com/youngwoo-yoon/youtube-gesture-dataset/blob/master/LICENSE"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Gesture Generation","url":"/task/gesture-generation","datasets_with_task":"/datasets/task/gesture-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["TED Gesture Dataset"],"data_loaders":[],"num_papers_in_archive":12,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/gesture-generation-on-ted-gesture-dataset","task":"Gesture Generation","dataset_variant":"TED Gesture Dataset","rows":6,"metrics":["FGD"],"first_row_in_archive_order":{"model":"AQ-GT","paper":"/paper/aq-gt-a-temporally-aligned-and-quantized-gru","metrics":{"FGD":"1.612"},"code_links":[{"title":"hvoss-techfak/AQGT","url":"https://github.com/hvoss-techfak/AQGT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/aq-gt-a-temporally-aligned-and-quantized-gru","title":"AQ-GT: a Temporally Aligned and Quantized GRU-Transformer for Co-Speech Gesture Synthesis","date":"2023-05-02","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/rhythmic-gesticulator-rhythm-aware-co-speech","title":"Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings","date":"2022-10-04","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/learning-hierarchical-cross-modal-association","title":"Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation","date":"2022-03-24","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/seeg-semantic-energized-co-speech-gesture","title":"SEEG: Semantic Energized Co-Speech Gesture Generation","date":"2022-01-01","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/speech2affectivegestures-synthesizing-co","title":"Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning","date":"2021-07-31","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/speech-gesture-generation-from-the-trimodal","title":"Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity","date":"2020-09-04","rows_on_this_dataset":1,"code_links":2,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}