{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-learning-for-video-classification-and","title":"Deep Learning for Video Classification and Captioning","arxiv_id":"1609.06782","date":"2016-09-22","proceeding":null,"authors":["Zuxuan Wu","Ting Yao","Yanwei Fu","Yu-Gang Jiang"],"abstract":"Accelerated by the tremendous increase in Internet bandwidth and storage\nspace, video data has been generated, published and spread explosively,\nbecoming an indispensable part of today's big data. In this paper, we focus on\nreviewing two lines of research aiming to stimulate the comprehension of videos\nwith deep learning: video classification and video captioning. While video\nclassification concentrates on automatically labeling video clips based on\ntheir semantic contents like human actions or complex events, video captioning\nattempts to generate a complete and natural sentence, enriching the single\nlabel as in video classification, to capture the most informative dynamics in\nvideos. In addition, we also provide a review of popular benchmarks and\ncompetitions, which are critical for evaluating the technical progress of this\nvibrant field.","url_abs":"http://arxiv.org/abs/1609.06782v2","url_pdf":"http://arxiv.org/pdf/1609.06782v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-learning-for-video-classification-and","repo_url":"https://github.com/mfayk/cu_icar_image_captioning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1609.06782","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}