{"url":"/task/image-to-text-retrieval","name":"Image-to-Text Retrieval","slug":"image-to-text-retrieval","description_markdown":"**Image-text retrieval** is the process of retrieving relevant images based on textual descriptions or finding corresponding textual descriptions for a given image. This task is interdisciplinary, combining techniques from computer vision, and natural language processing. The primary challenge lies in bridging the *semantic gap* — the difference between how visual data is represented in images and how humans describe that information using language. To address this, many methods focus on learning a shared embedding space where both images and text can be represented in a comparable way, allowing their similarities to be measured and facilitating more accurate retrieval.\r\n\r\n**Source**: <span class=\"description-source\"> [Extending CLIP for Category-to-Image Retrieval in E-commerce](https://arxiv.org/abs/2112.11294)</span>","categories":[{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":59,"papers_with_code":37,"benchmarks":8,"benchmark_tables_in_archive":8,"benchmark_tables_shown":8,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":8,"subtasks":0,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/image-to-text-retrieval-on-flickr30k","slug":"image-to-text-retrieval-on-flickr30k","dataset":"Flickr30k","dataset_url":"/dataset/flickr30k","rows_in_archive":11,"metrics":["Recall@1","Recall@5","Recall@10","Recall@Sum"],"first_row_in_archive_order":{"model":"InternVL-G-FT (finetuned, w/o ranking)","paper_title":"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks","paper_url":"/paper/internvl-scaling-up-vision-foundation-models","paper_date":"2023-12-21","arxiv_id":"2312.14238","code_links":[{"title":"opengvlab/internvl","url":"https://github.com/opengvlab/internvl"},{"title":"opengvlab/internvl-mmdetseg","url":"https://github.com/opengvlab/internvl-mmdetseg"}],"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}}},{"leaderboard":"/sota/image-to-text-retrieval-on-coco","slug":"image-to-text-retrieval-on-coco","dataset":"COCO (Common Objects in Context)","dataset_url":"/dataset/coco","rows_in_archive":9,"metrics":["Recall@1","Recall@5","Recall@10"],"first_row_in_archive_order":{"model":"BLIP-2 (ViT-G, fine-tuned)","paper_title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models","paper_url":"/paper/blip-2-bootstrapping-language-image-pre","paper_date":"2023-01-30","arxiv_id":"2301.12597","code_links":[{"title":"huggingface/transformers","url":"https://github.com/huggingface/transformers"},{"title":"salesforce/lavis","url":"https://github.com/salesforce/lavis"},{"title":"thudm/visualglm-6b","url":"https://github.com/thudm/visualglm-6b"},{"title":"baaivision/eva","url":"https://github.com/baaivision/eva"},{"title":"junshutang/Make-It-3D","url":"https://github.com/junshutang/Make-It-3D"},{"title":"facebookresearch/multimodal","url":"https://github.com/facebookresearch/multimodal"},{"title":"unispac/visual-adversarial-examples-jailbreak-large-language-models","url":"https://github.com/unispac/visual-adversarial-examples-jailbreak-large-language-models"},{"title":"yukw777/videoblip","url":"https://github.com/yukw777/videoblip"},{"title":"alibaba/graphtranslator","url":"https://github.com/alibaba/graphtranslator"},{"title":"gregor-ge/mblip","url":"https://github.com/gregor-ge/mblip"},{"title":"linzhiqiu/clip-flant5","url":"https://github.com/linzhiqiu/clip-flant5"},{"title":"kdr/videorag-mrr2024","url":"https://github.com/kdr/videorag-mrr2024"},{"title":"rabiulcste/vqazero","url":"https://github.com/rabiulcste/vqazero"},{"title":"jiwanchung/vlis","url":"https://github.com/jiwanchung/vlis"},{"title":"yangyucheng000/University","url":"https://github.com/yangyucheng000/University/tree/main/model-2/blip_2"},{"title":"2024-MindSpore-1/Code2","url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/blip_2"},{"title":"albertotestoni/ndq_visual_objects","url":"https://github.com/albertotestoni/ndq_visual_objects"}],"syntology":{"n":8,"n_ran":4,"n_unverified":4,"n_pointer_only":1}}},{"leaderboard":"/sota/image-to-text-retrieval-on-whoops","slug":"image-to-text-retrieval-on-whoops","dataset":"WHOOPS!","dataset_url":"/dataset/whoops","rows_in_archive":7,"metrics":["Specificity"],"first_row_in_archive_order":{"model":"BLIP2 FlanT5-XXL (Text-only FT)","paper_title":"Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images","paper_url":"/paper/breaking-common-sense-whoops-a-vision-and","paper_date":"2023-03-13","arxiv_id":"2303.07274","code_links":[],"syntology":null}},{"leaderboard":"/sota/image-to-text-retrieval-on-aic-icc","slug":"image-to-text-retrieval-on-aic-icc","dataset":"AIC-ICC","dataset_url":null,"rows_in_archive":2,"metrics":["Recall@1","Recall@10","Recall@5"],"first_row_in_archive_order":{"model":"ERNIE-ViL2.0","paper_title":"ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training","paper_url":"/paper/ernie-vil-2-0-multi-view-contrastive-learning","paper_date":"2022-09-30","arxiv_id":"2209.15270","code_links":[{"title":"PaddlePaddle/ERNIE","url":"https://github.com/PaddlePaddle/ERNIE"}],"syntology":null}},{"leaderboard":"/sota/image-to-text-retrieval-on-coco-1","slug":"image-to-text-retrieval-on-coco-1","dataset":"COCO","dataset_url":null,"rows_in_archive":1,"metrics":["Recall@1"],"first_row_in_archive_order":{"model":"SigLIP (ViT-L, zero-shot)","paper_title":"Sigmoid Loss for Language Image Pre-Training","paper_url":"/paper/sigmoid-loss-for-language-image-pre-training","paper_date":"2023-03-27","arxiv_id":"2303.15343","code_links":[{"title":"huggingface/transformers","url":"https://github.com/huggingface/transformers"},{"title":"mlfoundations/open_clip","url":"https://github.com/mlfoundations/open_clip"},{"title":"google-research/big_vision","url":"https://github.com/google-research/big_vision"},{"title":"apple/ml-mobileclip","url":"https://github.com/apple/ml-mobileclip"},{"title":"merveenoyan/siglip","url":"https://github.com/merveenoyan/siglip"},{"title":"borisdayma/clip-jax","url":"https://github.com/borisdayma/clip-jax"},{"title":"filipbasara0/simple-clip","url":"https://github.com/filipbasara0/simple-clip"},{"title":"morrisfl/unifex","url":"https://github.com/morrisfl/unifex"},{"title":"filipbasara0/relic","url":"https://github.com/filipbasara0/relic"},{"title":"shkarupa-alex/tfclip","url":"https://github.com/shkarupa-alex/tfclip"},{"title":"ramanakshay/clip","url":"https://github.com/ramanakshay/clip"}],"syntology":{"n":29,"n_ran":12,"n_unverified":17,"n_pointer_only":25}}},{"leaderboard":"/sota/image-to-text-retrieval-on-feta-car-manuals","slug":"image-to-text-retrieval-on-feta-car-manuals","dataset":"FETA Car-Manuals","dataset_url":"/dataset/feta-car-manuals","rows_in_archive":1,"metrics":["R@1","R@10","R@5"],"first_row_in_archive_order":{"model":"FETA's CLIP-MIL (Many-Shot Image-to-text)","paper_title":"FETA: Towards Specializing Foundation Models for Expert Task Applications","paper_url":"/paper/feta-towards-specializing-foundation-models","paper_date":"2022-09-08","arxiv_id":"2209.03648","code_links":[{"title":"alfassy/FETA","url":"https://github.com/alfassy/FETA"}],"syntology":null}},{"leaderboard":"/sota/image-to-text-retrieval-on-rsicd","slug":"image-to-text-retrieval-on-rsicd","dataset":"RSICD","dataset_url":"/dataset/rsicd","rows_in_archive":1,"metrics":["Image to Text Recall@1"],"first_row_in_archive_order":{"model":"GeoRSCLIP-FT","paper_title":"RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing","paper_url":"/paper/rs5m-a-large-scale-vision-language-dataset","paper_date":"2023-06-20","arxiv_id":"2306.11300","code_links":[{"title":"om-ai-lab/rs5m","url":"https://github.com/om-ai-lab/rs5m"}],"syntology":{"n":7,"n_ran":0,"n_unverified":7,"n_pointer_only":0}}},{"leaderboard":"/sota/image-to-text-retrieval-on-ruc-cas-wenlan","slug":"image-to-text-retrieval-on-ruc-cas-wenlan","dataset":"RUC-CAS-WenLan","dataset_url":null,"rows_in_archive":1,"metrics":["Recall@1","Recall@10","Recall@5"],"first_row_in_archive_order":{"model":"CMCL","paper_title":"WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training","paper_url":"/paper/wenlan-bridging-vision-and-language-by-large","paper_date":"2021-03-11","arxiv_id":"2103.06561","code_links":[{"title":"BAAI-WuDao/BriVl","url":"https://github.com/BAAI-WuDao/BriVl"},{"title":"Aman-4-Real/MMTG","url":"https://github.com/Aman-4-Real/MMTG"}],"syntology":null}}],"datasets":[{"url":"/dataset/coco","name":"COCO (Common Objects in Context)","full_name":"Common Objects in Context","num_papers_in_archive":11922},{"url":"/dataset/flickr30k","name":"Flickr30k","full_name":"Flickr30k","num_papers_in_archive":880},{"url":"/dataset/laion-400m","name":"LAION-400M","full_name":"","num_papers_in_archive":169},{"url":"/dataset/rsicd","name":"RSICD","full_name":"Remote Sensing Image Captioning Dataset","num_papers_in_archive":70},{"url":"/dataset/xm-3600","name":"XM 3600","full_name":"Crossmodal 3600","num_papers_in_archive":28},{"url":"/dataset/whoops","name":"WHOOPS!","full_name":"","num_papers_in_archive":14},{"url":"/dataset/fewsol","name":"FewSOL","full_name":"A Dataset for Few-Shot Object Learning in Robotic Environments","num_papers_in_archive":4},{"url":"/dataset/feta-car-manuals","name":"FETA Car-Manuals","full_name":"FETA Car-Manuals dataset, image-text retrieval for foundation models' expert data performance.","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":37,"tagged_in_all":59,"items":[{"url":"/paper/learning-transferable-visual-models-from","title":"Learning Transferable Visual Models From Natural Language Supervision","date":"2021-02-26","arxiv_id":"2103.00020","repositories_listed":82,"syntology":{"n":20,"n_ran":16,"n_unverified":4,"n_pointer_only":16}},{"url":"/paper/blip-2-bootstrapping-language-image-pre","title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models","date":"2023-01-30","arxiv_id":"2301.12597","repositories_listed":17,"syntology":{"n":8,"n_ran":4,"n_unverified":4,"n_pointer_only":1}},{"url":"/paper/sigmoid-loss-for-language-image-pre-training","title":"Sigmoid Loss for Language Image Pre-Training","date":"2023-03-27","arxiv_id":"2303.15343","repositories_listed":11,"syntology":{"n":29,"n_ran":12,"n_unverified":17,"n_pointer_only":25}},{"url":"/paper/align-before-fuse-vision-and-language","title":"Align before Fuse: Vision and Language Representation Learning with Momentum Distillation","date":"2021-07-16","arxiv_id":"2107.07651","repositories_listed":6,"syntology":{"n":5,"n_ran":3,"n_unverified":2,"n_pointer_only":3}},{"url":"/paper/flava-a-foundational-language-and-vision","title":"FLAVA: A Foundational Language And Vision Alignment Model","date":"2021-12-08","arxiv_id":"2112.04482","repositories_listed":4,"syntology":null},{"url":"/paper/oscar-object-semantics-aligned-pre-training","title":"Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks","date":"2020-04-13","arxiv_id":"2004.06165","repositories_listed":4,"syntology":{"n":23,"n_ran":11,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/deep-visual-semantic-alignments-for","title":"Deep Visual-Semantic Alignments for Generating Image Descriptions","date":"2014-12-07","arxiv_id":"1412.2306","repositories_listed":4,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/iglue-a-benchmark-for-transfer-learning","title":"IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and Languages","date":"2022-01-27","arxiv_id":"2201.11732","repositories_listed":3,"syntology":{"n":21,"n_ran":8,"n_unverified":13,"n_pointer_only":0}},{"url":"/paper/internvl-scaling-up-vision-foundation-models","title":"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks","date":"2023-12-21","arxiv_id":"2312.14238","repositories_listed":2,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/negative-pre-aware-for-noisy-cross-modal","title":"Negative Pre-aware for Noisy Cross-modal Matching","date":"2023-12-10","arxiv_id":"2312.05777","repositories_listed":2,"syntology":null},{"url":"/paper/multimodal-dataset-distillation-for-image","title":"Vision-Language Dataset Distillation","date":"2023-08-15","arxiv_id":"2308.07545","repositories_listed":2,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/one-peace-exploring-one-general","title":"ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities","date":"2023-05-18","arxiv_id":"2305.11172","repositories_listed":2,"syntology":{"n":7,"n_ran":2,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/upop-unified-and-progressive-pruning-for","title":"UPop: Unified and Progressive Pruning for Compressing Vision-Language Transformers","date":"2023-01-31","arxiv_id":"2301.13741","repositories_listed":2,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/hada-a-graph-based-amalgamation-framework-in","title":"HADA: A Graph-based Amalgamation Framework in Image-text Retrieval","date":"2023-01-11","arxiv_id":"2301.04742","repositories_listed":2,"syntology":null},{"url":"/paper/a-differentiable-semantic-metric","title":"A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal Retrieval","date":"2022-12-06","arxiv_id":null,"repositories_listed":2,"syntology":null},{"url":"/paper/altclip-altering-the-language-encoder-in-clip","title":"AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities","date":"2022-11-12","arxiv_id":"2211.06679","repositories_listed":2,"syntology":{"n":11,"n_ran":2,"n_unverified":9,"n_pointer_only":0}},{"url":"/paper/opt-omni-perception-pre-trainer-for-cross","title":"OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation","date":"2021-07-01","arxiv_id":"2107.00249","repositories_listed":2,"syntology":null},{"url":"/paper/wenlan-bridging-vision-and-language-by-large","title":"WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training","date":"2021-03-11","arxiv_id":"2103.06561","repositories_listed":2,"syntology":null},{"url":"/paper/exploring-models-and-data-for-remote-sensing","title":"Exploring Models and Data for Remote Sensing Image Caption Generation","date":"2017-12-21","arxiv_id":"1712.07835","repositories_listed":2,"syntology":null},{"url":"/paper/efficient-medical-vision-language-alignment","title":"Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models","date":"2025-06-10","arxiv_id":"2506.08990","repositories_listed":1,"syntology":null},{"url":"/paper/gabinsight-exploring-gender-activity-binding","title":"GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models","date":"2024-07-30","arxiv_id":"2407.21001","repositories_listed":1,"syntology":null},{"url":"/paper/towards-a-text-based-quantitative-and","title":"Towards a text-based quantitative and explainable histopathology image analysis","date":"2024-07-10","arxiv_id":"2407.07360","repositories_listed":1,"syntology":null},{"url":"/paper/bivlc-extending-vision-language","title":"BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval","date":"2024-06-14","arxiv_id":"2406.09952","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/efficient-remote-sensing-with-harmonized","title":"Efficient Remote Sensing with Harmonized Transfer Learning and Modality Alignment","date":"2024-04-28","arxiv_id":"2404.18253","repositories_listed":1,"syntology":{"n":6,"n_ran":1,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/linguistic-aware-patch-slimming-framework-for","title":"Linguistic-Aware Patch Slimming Framework for Fine-grained Cross-Modal Alignment","date":"2024-01-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/prototype-based-aleatoric-uncertainty-1","title":"Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval","date":"2023-09-29","arxiv_id":"2309.17093","repositories_listed":1,"syntology":{"n":19,"n_ran":12,"n_unverified":7,"n_pointer_only":0}},{"url":"/paper/prior-prototype-representation-joint-learning","title":"PRIOR: Prototype Representation Joint Learning from Medical Images and Reports","date":"2023-07-24","arxiv_id":"2307.12577","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/rs5m-a-large-scale-vision-language-dataset","title":"RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing","date":"2023-06-20","arxiv_id":"2306.11300","repositories_listed":1,"syntology":{"n":7,"n_ran":0,"n_unverified":7,"n_pointer_only":0}},{"url":"/paper/crossget-cross-guided-ensemble-of-tokens-for","title":"CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers","date":"2023-05-27","arxiv_id":"2305.17455","repositories_listed":1,"syntology":{"n":4,"n_ran":2,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/rethinking-benchmarks-for-cross-modal-image","title":"Rethinking Benchmarks for Cross-modal Image-text Retrieval","date":"2023-04-21","arxiv_id":"2304.10824","repositories_listed":1,"syntology":null}],"syntology_records":18,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}