{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/predicting-visual-features-from-text-for","title":"Predicting Visual Features from Text for Image and Video Caption Retrieval","arxiv_id":"1709.01362","date":"2017-09-05","proceeding":null,"authors":["Jianfeng Dong","Xirong Li","Cees G. M. Snoek"],"abstract":"This paper strives to find amidst a set of sentences the one best describing\nthe content of a given image or video. Different from existing works, which\nrely on a joint subspace for their image and video caption retrieval, we\npropose to do so in a visual space exclusively. Apart from this conceptual\nnovelty, we contribute \\emph{Word2VisualVec}, a deep neural network\narchitecture that learns to predict a visual feature representation from\ntextual input. Example captions are encoded into a textual embedding based on\nmulti-scale sentence vectorization and further transferred into a deep visual\nfeature of choice via a simple multi-layer perceptron. We further generalize\nWord2VisualVec for video caption retrieval, by predicting from text both 3-D\nconvolutional neural network features as well as a visual-audio representation.\nExperiments on Flickr8k, Flickr30k, the Microsoft Video Description dataset and\nthe very recent NIST TrecVid challenge for video caption retrieval detail\nWord2VisualVec's properties, its benefit over textual embeddings, the potential\nfor multimodal query composition and its state-of-the-art results.","url_abs":"http://arxiv.org/abs/1709.01362v3","url_pdf":"http://arxiv.org/pdf/1709.01362v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"predicting-visual-features-from-text-for","repo_url":"https://github.com/danieljf24/w2vv","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"video-description","task_name":"Video Description"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.01362","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}