{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/what-value-do-explicit-high-level-concepts","title":"What value do explicit high level concepts have in vision to language problems?","arxiv_id":"1506.01144","date":"2015-06-03","proceeding":"CVPR 2016 6","authors":["Qi Wu","Chunhua Shen","Lingqiao Liu","Anthony Dick","Anton Van Den Hengel"],"abstract":"Much of the recent progress in Vision-to-Language (V2L) problems has been\nachieved through a combination of Convolutional Neural Networks (CNNs) and\nRecurrent Neural Networks (RNNs). This approach does not explicitly represent\nhigh-level semantic concepts, but rather seeks to progress directly from image\nfeatures to text. We propose here a method of incorporating high-level concepts\ninto the very successful CNN-RNN approach, and show that it achieves a\nsignificant improvement on the state-of-the-art performance in both image\ncaptioning and visual question answering. We also show that the same mechanism\ncan be used to introduce external semantic information and that doing so\nfurther improves performance. In doing so we provide an analysis of the value\nof high level semantic information in V2L problems.","url_abs":"http://arxiv.org/abs/1506.01144v6","url_pdf":"http://arxiv.org/pdf/1506.01144v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"what-value-do-explicit-high-level-concepts","repo_url":"https://github.com/liuqihan/Image-Caption","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1506.01144","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}