{"url":"/method/altdiffusion","slug":"altdiffusion","name":"AltDiffusion","full_name":"AltDiffusion","full_name_withheld":false,"description_markdown":"In this work, we present a conceptually simple and effective method to train a strong bilingual multimodal representation model. Starting from the pretrained multimodal representation model CLIP released by OpenAI, we switched its text encoder with a pretrained multilingual text encoder XLM-R, and aligned both languages and image representations by a two-stage training schema consisting of teacher learning and contrastive learning. We validate our method through evaluations of a wide range of tasks. We set new state-of-the-art performances on a bunch of tasks including ImageNet-CN, Flicker30k- CN, and COCO-CN. Further, we obtain very close performances with CLIP on almost all tasks, suggesting that one can simply alter the text encoder in CLIP for extended capabilities such as multilingual understanding. Our models and code are available at https://github.com/FlagAI-Open/FlagAI.","description_state":"present","introduced_year":null,"introduced_by":{"title":"AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities","paper":"/paper/altclip-altering-the-language-encoder-in-clip","first_author":"Zhongzhi Chen","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/altclip-altering-the-language-encoder-in-clip"},"source":{"url":"https://arxiv.org/abs/2211.06679v2","title":"AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Generation Models","url":"/methods/category/image-generation-models","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":2,"papers_newest_first":[{"paper":"/paper/altdiffusion-a-multilingual-text-to-image","title":"AltDiffusion: A Multilingual Text-to-Image Diffusion Model","date":"2023-08-19","arxiv_id":"2308.09991","n_code_links":1,"syntology":{"ran":9,"of":11,"unverified":2,"pointer_only":11}},{"paper":"/paper/altclip-altering-the-language-encoder-in-clip","title":"AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities","date":"2022-11-12","arxiv_id":"2211.06679","n_code_links":2,"syntology":{"ran":2,"of":11,"unverified":9,"pointer_only":0}}],"papers_shown":2,"tasks":[{"task":"/task/blocking","name":"Blocking","papers":1},{"task":"/task/concept-alignment","name":"Concept Alignment","papers":1},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/cross-modal-retrieval","name":"Cross-Modal Retrieval","papers":1},{"task":"/task/image-classification","name":"Image Classification","papers":1},{"task":"/task/image-retrieval","name":"Image Retrieval","papers":1},{"task":"/task/image-to-text-retrieval","name":"Image-to-Text Retrieval","papers":1},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":1},{"task":"/task/text-to-image-generation","name":"Text-to-Image Generation","papers":1},{"task":"/task/xlm-r","name":"XLM-R","papers":1},{"task":"/task/zero-shot-cross-modal-retrieval","name":"Zero-Shot Cross-Modal Retrieval","papers":1},{"task":"/task/zero-shot-image-classification","name":"Zero-Shot Image Classification","papers":1},{"task":"/task/zero-shot-transfer-image-classification","name":"Zero-Shot Transfer Image Classification","papers":1},{"task":"/task/zero-shot-transfer-image-classification-cn","name":"Zero-Shot Transfer Image Classification (CN)","papers":1},{"task":"/task/zero-shot-image-retrieval","name":"Zero-shot Image Retrieval","papers":1},{"task":"/task/zero-shot-text-retrieval","name":"Zero-shot Text Retrieval","papers":1},{"task":"/task/model","name":"model","papers":1}],"tasks_shown":17,"n_tasks":17,"usage_by_year":[{"year":"2022","papers":1},{"year":"2023","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/altdiffusion"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}