{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semantic-speech-retrieval-with-a-visually","title":"Semantic speech retrieval with a visually grounded model of untranscribed speech","arxiv_id":"1710.01949","date":"2017-10-05","proceeding":null,"authors":["Herman Kamper","Gregory Shakhnarovich","Karen Livescu"],"abstract":"There is growing interest in models that can learn from unlabelled speech\npaired with visual context. This setting is relevant for low-resource speech\nprocessing, robotics, and human language acquisition research. Here we study\nhow a visually grounded speech model, trained on images of scenes paired with\nspoken captions, captures aspects of semantics. We use an external image tagger\nto generate soft text labels from images, which serve as targets for a neural\nmodel that maps untranscribed speech to (semantic) keyword labels. We introduce\na newly collected data set of human semantic relevance judgements and an\nassociated task, semantic speech retrieval, where the goal is to search for\nspoken utterances that are semantically relevant to a given text query. Without\nseeing any text, the model trained on parallel speech and images achieves a\nprecision of almost 60% on its top ten semantic retrievals. Compared to a\nsupervised model trained on transcriptions, our model matches human judgements\nbetter by some measures, especially in retrieving non-verbatim semantic\nmatches. We perform an extensive analysis of the model and its resulting\nrepresentations.","url_abs":"http://arxiv.org/abs/1710.01949v2","url_pdf":"http://arxiv.org/pdf/1710.01949v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"semantic-speech-retrieval-with-a-visually","repo_url":"https://github.com/kamperh/recipe_semantic_flickraudio","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"semantic-speech-retrieval-with-a-visually","repo_url":"https://github.com/kamperh/semantic_flickraudio","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"language-acquisition","task_name":"Language Acquisition"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1710.01949","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}