Papers › Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization

Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization

30 Jun 2018NeurIPS 2018 12arXiv:1807.00230archive 2025-07-28

Bruno Korbar, Du Tran, Lorenzo Torresani

There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio and video analysis from self-supervised temporal synchronization. We demonstrate that a calibrated curriculum learning scheme, a careful choice of negative examples, and the use of a contrastive loss are critical ingredients to obtain powerful multi-sensory representations from models optimized to discern temporal synchronization of audio-video pairs. Without further finetuning, the resulting audio features achieve performance superior or comparable to the state-of-the-art on established audio classification benchmarks (DCASE2014 and ESC-50). At the same time, our visual subnet provides a very effective initialization to improve the accuracy of video-based action recognition models: compared to learning from scratch, our self-supervised pretraining yields a remarkable gain of +19.9% in action recognition accuracy on UCF101 and a boost of +17.7% on HMDB51.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionAudio ClassificationSelf-Supervised Action RecognitionSelf-Supervised Audio ClassificationTemporal Action Localization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Classification ESC-50 AVTS Top-1 Accuracy 82.3 #28 of 29 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) AVTS Top-1 Accuracy 61.6 #10 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) AVTS 3-fold Accuracy 89.0 #10 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections