{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sequence-to-sequence-video-to-text","title":"Sequence to Sequence -- Video to Text","arxiv_id":"1505.00487","date":"2015-05-03","proceeding":null,"authors":["Subhashini Venugopalan","Marcus Rohrbach","Jeff Donahue","Raymond Mooney","Trevor Darrell","Kate Saenko"],"abstract":"Real-world videos often have complex dynamics; and methods for generating\nopen-domain video descriptions should be sensitive to temporal structure and\nallow both input (sequence of frames) and output (sequence of words) of\nvariable length. To approach this problem, we propose a novel end-to-end\nsequence-to-sequence model to generate captions for videos. For this we exploit\nrecurrent neural networks, specifically LSTMs, which have demonstrated\nstate-of-the-art performance in image caption generation. Our LSTM model is\ntrained on video-sentence pairs and learns to associate a sequence of video\nframes to a sequence of words in order to generate a description of the event\nin the video clip. Our model naturally is able to learn the temporal structure\nof the sequence of frames as well as the sequence model of the generated\nsentences, i.e. a language model. We evaluate several variants of our model\nthat exploit different visual features on a standard set of YouTube videos and\ntwo movie description datasets (M-VAD and MPII-MD).","url_abs":"http://arxiv.org/abs/1505.00487v3","url_pdf":"http://arxiv.org/pdf/1505.00487v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sequence-to-sequence-video-to-text","repo_url":"https://github.com/Kamino666/S2VT-video-caption","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"sequence-to-sequence-video-to-text","repo_url":"https://github.com/oddguan/audio_visual_video_caption","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"sequence-to-sequence-video-to-text","repo_url":"https://github.com/stillarrow/S2VT_ACT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"sequence-to-sequence-video-to-text","repo_url":"https://github.com/nasib-ullah/video-captioning-models-in-Pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"caption-generation","task_name":"Caption Generation"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1505.00487","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}