Browse State-of-the-Art › audio-visual learning
audio-visual learning
24 papers with code · 0 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
24 shown of 24 papers with code (38 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 May 2023 2 repositories listed Syntology ran 4 of 11 samples · 7 unverifiedIn this paper, we propose a novel method utilizing latent diffusion models trained for text-to-image-generation to generate images conditioned on audio recordings.
-
2 May 2025 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedSecond, we resolve conflicting optimization goals by separating contrastive and reconstruction objectives through dedicated global tokens.
-
1 Jan 2025 1 repository listedLong-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination.
-
17 Dec 2024 1 repository listed Syntology ran 0 of 14 samples · 14 unverified · 14 pointer-only (licence)Specifically, the CMCC module contains two branches: a cross-modal interaction branch and a temporal consistency-gated branch.
-
29 Aug 2024 1 repository listedTo address this issue, we propose a novel audio-visual learning framework which is instantiated with two individual learning schemes: self-supervised predictive learning (SSPL) and semantic-aware contrastive learning…
-
7 Jun 2024 1 repository listedFurthermore, to suppress the background features in each modality from foreground matched audio-visual features, we introduce a robust discriminative foreground mining scheme.
-
14 Mar 2024 1 repository listed Syntology ran 6 of 13 samples · 7 unverifiedRecent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations.
-
29 Nov 2023 1 repository listedThe prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA) towards SVs.
-
21 Nov 2023 1 repository listedRecent methods mainly focus on learning multi-modal features aligned with class names to enhance the generalization ability to unseen categories.
-
7 Nov 2023 1 repository listed Syntology ran 10 of 10 samples · 0 unverified · 10 pointer-only (licence)Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment.
-
11 Oct 2023 1 repository listedTo implement the prior knowledge, we first train the audio-visual network, which learns the correspondence between auditory and visual information.
-
19 Sep 2023 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information.
-
11 Sep 2023 1 repository listed Syntology ran 6 of 8 samples · 2 unverifiedOur CIGN leverages learnable audio-visual class tokens and audio-visual grouping to continually aggregate class-aware features.
-
30 May 2023 1 repository listedThe ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task.
-
27 May 2023 1 repository listed Syntology ran 6 of 6 samples · 0 unverifiedAudio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its modality-aligned setting, i.
-
6 Apr 2023 1 repository listedWe show empirical results that demonstrate the effectiveness of our benchmark.
-
7 Feb 2023 1 repository listedSpecifically, we explore the effects of pre-trained models on two audio-visual learning scenarios: cross-modal initialization and multi-modal joint learning.
-
29 Jul 2022 1 repository listed Syntology ran 6 of 14 samples · 8 unverifiedConventional audio-visual models have independent audio and video branches.
-
12 Jul 2022 1 repository listed Syntology ran 2 of 10 samples · 8 unverifiedIn this paper, we analyze the modality asynchrony and undifferentiated instances phenomena of the multiple instance learning (MIL) procedure, and further investigate its negative impact on weakly-supervised audio-visual…
-
26 Mar 2022 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedIn this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos.
-
8 Nov 2021 1 repository listedIn this paper, we explore self-supervised audio-visual models that learn from instructional videos.
-
22 Apr 2021 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedHaving access to multi-modal cues (e.
-
5 Apr 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we propose to make a systematic study on machines multisensory perception under attacks.
-
12 Jan 2021 1 repository listedAML aims to generate a modality-independent representation for each person in each modality via adversarial learning, while simultaneously learns a robust similarity measure for cross-modality matching via metric…
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections