{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/frame-and-segment-level-features-and","title":"Frame- and Segment-Level Features and Candidate Pool Evaluation for Video Caption Generation","arxiv_id":"1608.04959","date":"2016-08-17","proceeding":null,"authors":["Rakshith Shetty","Jorma Laaksonen"],"abstract":"We present our submission to the Microsoft Video to Language Challenge of\ngenerating short captions describing videos in the challenge dataset. Our model\nis based on the encoder--decoder pipeline, popular in image and video\ncaptioning systems. We propose to utilize two different kinds of video\nfeatures, one to capture the video content in terms of objects and attributes,\nand the other to capture the motion and action information. Using these diverse\nfeatures we train models specializing in two separate input sub-domains. We\nthen train an evaluator model which is used to pick the best caption from the\npool of candidates generated by these domain expert models. We argue that this\napproach is better suited for the current video captioning task, compared to\nusing a single model, due to the diversity in the dataset.\n  Efficacy of our method is proven by the fact that it was rated best in MSR\nVideo to Language Challenge, as per human evaluation. Additionally, we were\nranked second in the automatic evaluation metrics based table.","url_abs":"http://arxiv.org/abs/1608.04959v1","url_pdf":"http://arxiv.org/pdf/1608.04959v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"frame-and-segment-level-features-and","repo_url":"https://github.com/rakshithShetty/captionGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"caption-generation","task_name":"Caption Generation"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"video-captioning","task_name":"Video Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}