{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-semantic-sentence-representations","title":"Learning semantic sentence representations from visually grounded language without lexical knowledge","arxiv_id":"1903.11393","date":"2019-03-27","proceeding":null,"authors":["Danny Merkx","Stefan Frank"],"abstract":"Current approaches to learning semantic representations of sentences often\nuse prior word-level knowledge. The current study aims to leverage visual\ninformation in order to capture sentence level semantics without the need for\nword embeddings. We use a multimodal sentence encoder trained on a corpus of\nimages with matching text captions to produce visually grounded sentence\nembeddings. Deep Neural Networks are trained to map the two modalities to a\ncommon embedding space such that for an image the corresponding caption can be\nretrieved and vice versa. We show that our model achieves results comparable to\nthe current state-of-the-art on two popular image-caption retrieval benchmark\ndata sets: MSCOCO and Flickr8k. We evaluate the semantic content of the\nresulting sentence embeddings using the data from the Semantic Textual\nSimilarity benchmark task and show that the multimodal embeddings correlate\nwell with human semantic similarity judgements. The system achieves\nstate-of-the-art results on several of these benchmarks, which shows that a\nsystem trained solely on multimodal data, without assuming any word\nrepresentations, is able to capture sentence level semantics. Importantly, this\nresult shows that we do not need prior knowledge of lexical level semantics in\norder to model sentence level semantics. These findings demonstrate the\nimportance of visual information in semantics.","url_abs":"http://arxiv.org/abs/1903.11393v1","url_pdf":"http://arxiv.org/pdf/1903.11393v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-semantic-sentence-representations","repo_url":"https://github.com/DannyMerkx/caption2image","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"grounded-language-learning","task_name":"Grounded language learning"},{"task_slug":"learning-semantic-representations","task_name":"Learning Semantic Representations"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentence-embeddings","task_name":"Sentence Embeddings"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}