{"url":"/method/camoe","slug":"camoe","name":"CAMoE","full_name":"CAMoE","full_name_withheld":false,"description_markdown":"**CAMoE** is a multi-stream Corpus Alignment network with single gate Mixture-of-Experts (MoE) for video-text retrieval. The CAMoE employs Mixture-of-Experts (MoE) to extract multi-perspective video representations, including action, entity, scene, etc., then align them with the corresponding part of the text. A [Dual Softmax Loss](https://paperswithcode.com/method/dual-softmax-loss) (DSL) is used to avoid the one-way optimum-match which occurs in previous contrastive methods. Introducing the intrinsic prior of each pair in a batch, DSL serves as a reviser to correct the similarity matrix and achieves the dual optimal match.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss","paper":"/paper/improving-video-text-retrieval-by-multi","first_author":"Xing Cheng","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/improving-video-text-retrieval-by-multi"},"source":{"url":"https://arxiv.org/abs/2109.04290v3","title":"Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Video-Text Retrieval Models","url":"/methods/category/video-text-retrieval-models","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":2,"papers_newest_first":[{"paper":null,"title":"An Audio-centric Multi-task Learning Framework for Streaming Ads Targeting on Spotify","date":"2025-06-23","arxiv_id":"2506.18735","n_code_links":0,"syntology":null},{"paper":"/paper/improving-video-text-retrieval-by-multi","title":"Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss","date":"2021-09-09","arxiv_id":"2109.04290","n_code_links":2,"syntology":null}],"papers_shown":2,"tasks":[{"task":"/task/mixture-of-experts","name":"Mixture-of-Experts","papers":2},{"task":"/task/click-through-rate-prediction","name":"Click-Through Rate Prediction","papers":1},{"task":"/task/multi-task-learning","name":"Multi-Task Learning","papers":1},{"task":"/task/retrieval","name":"Retrieval","papers":1},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":1},{"task":"/task/video-retrieval","name":"Video Retrieval","papers":1},{"task":"/task/video-text-retrieval","name":"Video-Text Retrieval","papers":1}],"tasks_shown":7,"n_tasks":7,"usage_by_year":[{"year":"2021","papers":1},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/camoe"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}