{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/using-descriptive-video-services-to-create-a","title":"Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research","arxiv_id":"1503.01070","date":"2015-03-03","proceeding":null,"authors":["Atousa Torabi","Christopher Pal","Hugo Larochelle","Aaron Courville"],"abstract":"In this work, we introduce a dataset of video annotated with high quality\nnatural language phrases describing the visual content in a given segment of\ntime. Our dataset is based on the Descriptive Video Service (DVS) that is now\nencoded on many digital media products such as DVDs. DVS is an audio narration\ndescribing the visual elements and actions in a movie for the visually\nimpaired. It is temporally aligned with the movie and mixed with the original\nmovie soundtrack. We describe an automatic DVS segmentation and alignment\nmethod for movies, that enables us to scale up the collection of a DVS-derived\ndataset with minimal human intervention. Using this method, we have collected\nthe largest DVS-derived dataset for video description of which we are aware.\nOur dataset currently includes over 84.6 hours of paired video/sentences from\n92 DVDs and is growing.","url_abs":"http://arxiv.org/abs/1503.01070v1","url_pdf":"http://arxiv.org/pdf/1503.01070v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"using-descriptive-video-services-to-create-a","repo_url":"https://github.com/jssprz/video_captioning_datasets","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"descriptive","task_name":"Descriptive"},{"task_slug":"video-description","task_name":"Video Description"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1503.01070","atlas_url":"https://app.syntology.ai/?focus=1503.01070","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}