{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/avhbench-a-cross-modal-hallucination","title":"AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models","arxiv_id":"2410.18325","date":"2024-10-23","proceeding":null,"authors":["Kim Sung-Bin","Oh Hyun-Bin","JungMok Lee","Arda Senocak","Joon Son Chung","Tae-Hyun Oh"],"abstract":"Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understanding of the world. In recognition of this fact, audio-visual LLMs have recently emerged. Despite promising developments, the lack of dedicated benchmarks poses challenges for understanding and evaluating models. In this work, we show that audio-visual LLMs struggle to discern subtle relationships between audio and visual signals, leading to hallucinations, underscoring the need for reliable benchmarks. To address this, we introduce AVHBench, the first comprehensive benchmark specifically designed to evaluate the perception and comprehension capabilities of audio-visual LLMs. Our benchmark includes tests for assessing hallucinations, as well as the cross-modal matching and reasoning abilities of these models. Our results reveal that most existing audio-visual LLMs struggle with hallucinations caused by cross-interactions between modalities, due to their limited capacity to perceive complex multimodal signals and their relationships. Additionally, we demonstrate that simple training with our AVHBench improves robustness of audio-visual LLMs against hallucinations.","url_abs":"https://arxiv.org/abs/2410.18325v1","url_pdf":"https://arxiv.org/pdf/2410.18325v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"hallucination","task_name":"Hallucination"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.18325","atlas_url":"https://app.syntology.ai/?focus=2410.18325","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.18325"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/kaist-ami/AVHBench","reach":null}],"summary":{"ran_draft_wrong":1,"unverified":2},"by_repo_kind":{"found_in_text":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"c4328006bd37dd35","entry":"func_gain_name2value","repo":"kaist-ami/AVHBench","repo_kind":"found_in_text","path":"AVHBench-Align-FT/evaluation.py","file_url":"https://github.com/kaist-ami/AVHBench/blob/HEAD/AVHBench-Align-FT/evaluation.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c4328006bd37dd35"}},{"code_sha256_prefix":"bfd33d4ca82a6500","entry":"gradio_answer","repo":"kaist-ami/AVHBench","repo_kind":"found_in_text","path":"AVHBench-Align-FT/inference.py","file_url":"https://github.com/kaist-ami/AVHBench/blob/HEAD/AVHBench-Align-FT/inference.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bfd33d4ca82a6500"}},{"code_sha256_prefix":"650feda504944cf1","entry":"gradio_ask","repo":"kaist-ami/AVHBench","repo_kind":"found_in_text","path":"AVHBench-Align-FT/inference.py","file_url":"https://github.com/kaist-ami/AVHBench/blob/HEAD/AVHBench-Align-FT/inference.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"650feda504944cf1"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}