Papers › Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening...

Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking

3 Nov 2025arXiv:2511.01299added by Syntology

Siyin Wang, Zengrui Jin, Changli Tang, Qiujia Li, Bo Li, Chen Chen, Yuchen Hu, Wenyi Yu, Yixuan Li, Jimin Zhuang, Yudong Yang, Mingqiu Wang, Michael Han, Yifan Ding, Junwen Bai, Tom Ouyang, Shuo-yiin Chang, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Guangzhi Sun, Zhehuai Chen, Ji Wu, Bowen Zhou, Yuxuan Wang, Tara Sainath, Yonghui Wu, Chao Zhang

Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).

In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities of foundation models, towards more comprehensive understanding, more natural generation and more human-like interaction. Audio, as a modality rich in semantic, emotional, and contextual cues, plays a vital role in achieving naturalistic and embodied machine intelligence. This survey provides a comprehensive review of recent progress in integrating audio into LLMs, with a focus on four key areas: audio comprehension, audio generation, speech-based interaction, and audio-visual understanding. We analyze how LLMs are reshaping audio perception and reasoning, enabling systems to understand sound at a deeper semantic level, generate expressive audio outputs, and engage in human-like spoken interaction. Furthermore, we explore how the fusion of audio and visual modalities enhances situational awareness and cross-modal reasoning, pushing the boundaries of multimodal intelligence. This survey not only synthesizes existing research but also identifies critical challenges and future directions for building audio-native AGI systems capable of perceiving, understanding, and interacting through sound as naturally as humans do.

PaperPDF

In Syntology View this paper on Syntology, its page in Syntology's graph. That page lists the repositories linked to the paper, the abstract and the calls for agents.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, on Syntology's MCP service (how to connect):

Code

suno-ai/bark found in paper text by SyntologyMITSyntology: no sample linked to this paper (harvested for another paper). Syntology's graph links these samples to Scaling Properties of Speech Language Models (arXiv:2404.00685; that paper's own run record: 0 ran · 3 unverified); nothing checks that they implement this paper's method. report

Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

A paper named beside a repository with no sample linked to this paper is shown with that paper's own run record, not this paper's: “ran” means executed on a synthesized input, not that the code is correct, and “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. Papers are listed in arXiv-id order, at most three per repository.

Code Syntology ran Syntology

Syntology holds the repository link but has not harvested or run code from it.

Results from the paper

The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2511.01299, was issued in November 2025, after that date, so the archive has no leaderboard rows for it.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections