Papers › Coordinated Joint Multimodal Embeddings for Generalized Audio-Visual Zeroshot...
Coordinated Joint Multimodal Embeddings for Generalized Audio-Visual Zeroshot Classification and Retrieval of Videos
Kranti Kumar Parida, Neeraj Matiyali, Tanaya Guha, Gaurav Sharma
We present an audio-visual multimodal approach for the task of zeroshot learning (ZSL) for classification and retrieval of videos. ZSL has been studied extensively in the recent past but has primarily been limited to visual modality and to images. We demonstrate that both audio and visual modalities are important for ZSL for videos. Since a dataset to study the task is currently not available, we also construct an appropriate multimodal dataset with 33 classes containing 156,416 videos, from an existing large scale audio event dataset. We empirically show that the performance improves by adding audio modality for both tasks of zeroshot classification and retrieval, when using multimodal extensions of embedding learning methods. We also propose a novel method to predict the `dominant' modality using a jointly learned modality attention network. We learn the attention in a semi-supervised setting and thus do not require any additional explicit labelling for the modalities. We provide qualitative validation of the modality specific attention, which also successfully generalizes to unseen test classes.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| GZSL Video Classification | ActivityNet-GZSL(main) | CJME | HM | 5.12 | #7 of 7 | Archive leaderboard | report |
| GZSL Video Classification | ActivityNet-GZSL(main) | CJME | ZSL | 5.84 | #7 of 7 | Archive leaderboard | report |
| GZSL Video Classification | UCF-GZSL(main) | CJME | HM | 12.48 | #7 of 7 | Archive leaderboard | report |
| GZSL Video Classification | UCF-GZSL(main) | CJME | ZSL | 8.29 | #7 of 7 | Archive leaderboard | report |
| GZSL Video Classification | VGGSound-GZSL(main) | CJME | HM | 6.17 | #5 of 7 | Archive leaderboard | report |
| GZSL Video Classification | VGGSound-GZSL(main) | CJME | ZSL | 5.16 | #5 of 7 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections