{"url":"/task/cross-modal-information-retrieval","name":"Cross-Modal Information Retrieval","slug":"cross-modal-information-retrieval","description_markdown":"**Cross-Modal Information Retrieval** (CMIR) is the task of finding relevant items across different modalities. For example, given an image, find a text or vice versa. The main challenge in CMIR is known as the *heterogeneity gap*: since items from different modalities have different data types, the similarity between them cannot be measured directly. Therefore, the majority of CMIR methods published to date attempt to bridge this gap by learning a latent representation space, where the similarity between items from different modalities can be measured.\r\n\r\n<span class=\"description-source\">Source: [Scene-centric vs. Object-centric Image-Text Cross-modal Retrieval: A Reproducibility Study](https://arxiv.org/abs/2301.05174)</span>","categories":[{"name":"Miscellaneous","url":"/area/miscellaneous"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":16,"papers_with_code":8,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":1,"subtasks":1,"parent_tasks":1},"benchmarks":[],"datasets":[{"url":"/dataset/ted-vcr","name":"TED VCR","full_name":"","num_papers_in_archive":1}],"subtasks":[{"url":"/task/cross-modal-retrieval","name":"Cross-Modal Retrieval"}],"parent_tasks":[{"url":"/task/multi-modal","name":"Image Retrieval with Multi-Modal Query"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":8,"of":8,"tagged_in_all":16,"items":[{"url":"/paper/picture-it-in-your-mind-generating-high-level","title":"Picture It In Your Mind: Generating High Level Visual Representations From Textual Descriptions","date":"2016-06-23","arxiv_id":"1606.07287","repositories_listed":2,"syntology":null},{"url":"/paper/lseh-semantically-enhanced-hard-negatives-for","title":"Improving Visual-Semantic Embeddings by Learning Semantically-Enhanced Hard Negatives for Cross-modal Information Retrieval","date":"2022-10-10","arxiv_id":"2210.04754","repositories_listed":1,"syntology":null},{"url":"/paper/visualsparta-sparse-transformer-fragment","title":"VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words","date":"2021-01-01","arxiv_id":"2101.00265","repositories_listed":1,"syntology":null},{"url":"/paper/learning-the-best-pooling-strategy-for-visual","title":"Learning the Best Pooling Strategy for Visual Semantic Embedding","date":"2020-11-09","arxiv_id":"2011.04305","repositories_listed":1,"syntology":null},{"url":"/paper/fine-grained-visual-textual-alignment-for","title":"Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders","date":"2020-08-12","arxiv_id":"2008.05231","repositories_listed":1,"syntology":{"n":16,"n_ran":2,"n_unverified":14,"n_pointer_only":0}},{"url":"/paper/zscrgan-a-gan-based-expectation-maximization","title":"ZSCRGAN: A GAN-based Expectation Maximization Model for Zero-Shot Retrieval of Images from Textual Descriptions","date":"2020-07-23","arxiv_id":"2007.12212","repositories_listed":1,"syntology":null},{"url":"/paper/approaching-small-molecule-prioritization-as","title":"Cross-modal representation alignment of molecular structure and perturbation-induced transcriptional profiles","date":"2019-11-22","arxiv_id":"1911.10241","repositories_listed":1,"syntology":null},{"url":"/paper/cmir-net-a-deep-learning-based-model-for","title":"CMIR-NET : A Deep Learning Based Model For Cross-Modal Retrieval In Remote Sensing","date":"2019-04-09","arxiv_id":"1904.04794","repositories_listed":1,"syntology":null}],"syntology_records":1,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}