{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tgif-a-new-dataset-and-benchmark-on-animated","title":"TGIF: A New Dataset and Benchmark on Animated GIF Description","arxiv_id":"1604.02748","date":"2016-04-10","proceeding":"CVPR 2016 6","authors":["Yuncheng Li","Yale Song","Liangliang Cao","Joel Tetreault","Larry Goldberg","Alejandro Jaimes","Jiebo Luo"],"abstract":"With the recent popularity of animated GIFs on social media, there is need\nfor ways to index them with rich metadata. To advance research on animated GIF\nunderstanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K\nanimated GIFs from Tumblr and 120K natural language descriptions obtained via\ncrowdsourcing. The motivation for this work is to develop a testbed for image\nsequence description systems, where the task is to generate natural language\ndescriptions for animated GIFs or video clips. To ensure a high quality\ndataset, we developed a series of novel quality controls to validate free-form\ntext input from crowdworkers. We show that there is unambiguous association\nbetween visual content and natural language descriptions in our dataset, making\nit an ideal benchmark for the visual content captioning task. We perform\nextensive statistical analyses to compare our dataset to existing image and\nvideo description datasets. Next, we provide baseline results on the animated\nGIF description task, using three representative techniques: nearest neighbor,\nstatistical machine translation, and recurrent neural networks. Finally, we\nshow that models fine-tuned from our animated GIF description dataset can be\nhelpful for automatic movie description.","url_abs":"http://arxiv.org/abs/1604.02748v2","url_pdf":"http://arxiv.org/pdf/1604.02748v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tgif-a-new-dataset-and-benchmark-on-animated","repo_url":"https://github.com/raingo/TGIF-Release","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"text-generation","task_name":"Text Generation"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"video-description","task_name":"Video Description"}],"methods":[],"datasets_introduced":[{"slug":"tgif","name":"TGIF","full_name":"Tumblr GIF"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1604.02748","atlas_url":"https://app.syntology.ai/?focus=1604.02748","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}