Papers › ViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction

ViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction

11 Dec 2020arXiv:2012.06170archive 2025-07-28

Samyak Jain, Pradeep Yarlagadda, Shreyank Jyoti, Shyamgopal Karthik, Ramanathan Subramanian, Vineet Gandhi

We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the decoder infers a saliency map via trilinear interpolation and 3D convolutions, combining features from multiple hierarchies. The overall architecture of ViNet is conceptually simple; it is causal and runs in real-time (60 fps). ViNet does not use audio as input and still outperforms the state-of-the-art audio-visual saliency prediction models on nine different datasets (three visual-only and six audio-visual datasets). ViNet also surpasses human performance on the CC, SIM and AUC metrics for the AVE dataset, and to our knowledge, it is the first network to do so. We also explore a variation of ViNet architecture by augmenting audio features into the decoder. To our surprise, upon sufficient training, the network becomes agnostic to the input audio and provides the same output irrespective of the input. Interestingly, we also observe similar behaviour in the previous state-of-the-art models \cite{tsiami2020stavis} for audio-visual saliency prediction. Our findings contrast with previous works on deep learning-based audio-visual saliency prediction, suggesting a clear avenue for future explorations incorporating audio in a more effective manner. The code and pre-trained models are available at https://github.com/samyak0210/ViNet.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

samyak0210/ViNet officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionDecoderPredictionSaliency PredictionVideo Saliency DetectionVideo Saliency Prediction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Saliency Detection DHF1K ViNet AUC-J 0.908 #1 of 3 Archive leaderboard report
Video Saliency Detection DHF1K ViNet CC 0.51 #1 of 3 Archive leaderboard report
Video Saliency Detection DHF1K ViNet NSS 2.87 #1 of 3 Archive leaderboard report
Video Saliency Detection DHF1K ViNet s-AUC 0.728 #1 of 3 Archive leaderboard report
Video Saliency Detection DIEM AViNet CC 0.632 #1 of 1 Archive leaderboard report
Video Saliency Detection Hollywood2 ViNet CC 0.693 #1 of 1 Archive leaderboard report
Video Saliency Detection MSU Video Saliency Prediction ViNet (dave) AUC-J 0.864 #1 of 14 Archive leaderboard report
Video Saliency Detection MSU Video Saliency Prediction ViNet (dave) CC 0.733 #1 of 14 Archive leaderboard report
Video Saliency Detection MSU Video Saliency Prediction ViNet (dave) FPS 1.10 #1 of 14 Archive leaderboard report
Video Saliency Detection MSU Video Saliency Prediction ViNet (dave) KLDiv 0.497 #1 of 14 Archive leaderboard report
Video Saliency Detection MSU Video Saliency Prediction ViNet (dave) NSS 2.13 #1 of 14 Archive leaderboard report
Video Saliency Detection MSU Video Saliency Prediction ViNet (dave) SIM 0.627 #1 of 14 Archive leaderboard report
Video Saliency Detection UCFSports ViNet CC 0.673 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections