{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-modal-retrieval-in-the-cooking-context","title":"Cross-Modal Retrieval in the Cooking Context: Learning Semantic Text-Image Embeddings","arxiv_id":"1804.11146","date":"2018-04-30","proceeding":null,"authors":["Micael Carvalho","Rémi Cadène","David Picard","Laure Soulier","Nicolas Thome","Matthieu Cord"],"abstract":"Designing powerful tools that support cooking activities has rapidly gained\npopularity due to the massive amounts of available data, as well as recent\nadvances in machine learning that are capable of analyzing them. In this paper,\nwe propose a cross-modal retrieval model aligning visual and textual data (like\npictures of dishes and their recipes) in a shared representation space. We\ndescribe an effective learning scheme, capable of tackling large-scale\nproblems, and validate it on the Recipe1M dataset containing nearly 1 million\npicture-recipe pairs. We show the effectiveness of our approach regarding\nprevious state-of-the-art models and present qualitative results over\ncomputational cooking use cases.","url_abs":"http://arxiv.org/abs/1804.11146v1","url_pdf":"http://arxiv.org/pdf/1804.11146v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cross-modal-retrieval-in-the-cooking-context","repo_url":"https://github.com/Cadene/recipe1m.bootstrap.pytorch","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-modal-retrieval-on-recipe1m","task":"Cross-Modal Retrieval","dataset":"Recipe1M","model":"AdaMine","rank_in_archive_order":9,"of":9,"metrics":{"Image-to-text R@1":"39.8","Text-to-image R@1":"40.2"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.11146","atlas_url":"https://app.syntology.ai/?focus=1804.11146","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}