Papers › Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text Retrieval

Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text Retrieval

11 Jun 2018ICMR 2018 6archive 2025-07-28

Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, Amit K. Roy-Chowdhury

Constructing a joint representation invariant across different modalities (e.g., video, language) is of significant importance in many multimedia applications. While there are a number of recent successes in developing effective image-text retrieval methods by learning joint representations, the video-text retrieval task, in contrast, has not been explored to its fullest extent. In this paper, we study how to effectively utilize available multi-modal cues from videos for the cross-modal video-text retrieval task. Based on our analysis, we propose a novel framework that simultaneously utilizes multimodal features (different visual characteristics, audio inputs, and text) by a fusion strategy for efficient retrieval. Furthermore, we explore several loss functions in training the joint embedding and propose a modified pairwise ranking loss for the retrieval task. Experiments on MSVD and MSR-VTT datasets demonstrate that our method achieves significant performance gain compared to the state-of-the-art approaches.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image-text RetrievalRetrievalText RetrievalVideo RetrievalVideo-Text Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Retrieval MSR-VTT JEMC text-to-video Mean Rank 213.8 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC text-to-video Median Rank 29.7 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC text-to-video R@1 7.0 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC text-to-video R@10 29.7 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC text-to-video R@5 20.9 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC video-to-text Mean Rank 134 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC video-to-text Median Rank 16 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC video-to-text R@1 12.5 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC video-to-text R@10 42.2 #38 of 40 Archive leaderboard report
Video Retrieval MSR-VTT JEMC video-to-text R@5 32.1 #38 of 40 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections