{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-language-visual-embedding-for-movie","title":"Learning Language-Visual Embedding for Movie Understanding with Natural-Language","arxiv_id":"1609.08124","date":"2016-09-26","proceeding":null,"authors":["Atousa Torabi","Niket Tandon","Leonid Sigal"],"abstract":"Learning a joint language-visual embedding has a number of very appealing\nproperties and can result in variety of practical application, including\nnatural language image/video annotation and search. In this work, we study\nthree different joint language-visual neural network model architectures. We\nevaluate our models on large scale LSMDC16 movie dataset for two tasks: 1)\nStandard Ranking for video annotation and retrieval 2) Our proposed movie\nmultiple-choice test. This test facilitate automatic evaluation of\nvisual-language models for natural language video annotation based on human\nactivities. In addition to original Audio Description (AD) captions, provided\nas part of LSMDC16, we collected and will make available a) manually generated\nre-phrasings of those captions obtained using Amazon MTurk b) automatically\ngenerated human activity elements in \"Predicate + Object\" (PO) phrases based on\n\"Knowlywood\", an activity knowledge mining model. Our best model archives\nRecall@10 of 19.2% on annotation and 18.9% on video retrieval tasks for subset\nof 1000 samples. For multiple-choice test, our best model achieve accuracy\n58.11% over whole LSMDC16 public test-set.","url_abs":"http://arxiv.org/abs/1609.08124v1","url_pdf":"http://arxiv.org/pdf/1609.08124v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-retrieval-on-msr-vtt","task":"Video Retrieval","dataset":"MSR-VTT","model":"C+LSTM+SA+FC7","rank_in_archive_order":40,"of":40,"metrics":{"text-to-video Median Rank":"55","text-to-video R@1":"4.2","text-to-video R@10":"19.9","video-to-text R@5":"12.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1609.08124","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}