{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-the-visual-feature-space-for","title":"Exploring The Visual Feature Space for Multimodal Neural Decoding","arxiv_id":"2505.15755","date":"2025-05-21","proceeding":null,"authors":["Weihao Xia","Cengiz Oztireli"],"abstract":"The intrication of brain signals drives research that leverages multimodal AI to align brain modalities with visual and textual data for explainable descriptions. However, most existing studies are limited to coarse interpretations, lacking essential details on object descriptions, locations, attributes, and their relationships. This leads to imprecise and ambiguous reconstructions when using such cues for visual decoding. To address this, we analyze different choices of vision feature spaces from pre-trained visual components within Multimodal Large Language Models (MLLMs) and introduce a zero-shot multimodal brain decoding method that interacts with these models to decode across multiple levels of granularities. % To assess a model's ability to decode fine details from brain signals, we propose the Multi-Granularity Brain Detail Understanding Benchmark (MG-BrainDub). This benchmark includes two key tasks: detailed descriptions and salient question-answering, with metrics highlighting key visual elements like objects, attributes, and relationships. Our approach enhances neural decoding precision and supports more accurate neuro-decoding applications. Code will be available at https://github.com/weihaox/VINDEX.","url_abs":"https://arxiv.org/abs/2505.15755v1","url_pdf":"https://arxiv.org/pdf/2505.15755v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"brain-decoding","task_name":"Brain Decoding"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[{"method_slug":"align","method_name":"ALIGN"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2505.15755","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.15755"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/weihaox/VINDEX","reach":null}],"summary":{"ran":4},"by_repo_kind":{"found_in_text":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"ea01971e26a7f81f","entry":"BrainX","repo":"weihaox/VINDEX","repo_kind":"found_in_text","path":"src/model/model.py","file_url":"https://github.com/weihaox/VINDEX/blob/HEAD/src/model/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ea01971e26a7f81f"}},{"code_sha256_prefix":"c3008658a1e2f570","entry":"Perceiver","repo":"weihaox/VINDEX","repo_kind":"found_in_text","path":"src/model/model.py","file_url":"https://github.com/weihaox/VINDEX/blob/HEAD/src/model/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c3008658a1e2f570"}},{"code_sha256_prefix":"11192fbd2327ed4c","entry":"PerceiverAttention","repo":"weihaox/VINDEX","repo_kind":"found_in_text","path":"src/model/model.py","file_url":"https://github.com/weihaox/VINDEX/blob/HEAD/src/model/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"11192fbd2327ed4c"}},{"code_sha256_prefix":"cddbdad94b8e10ad","entry":"PerceiverResampler","repo":"weihaox/VINDEX","repo_kind":"found_in_text","path":"src/model/model.py","file_url":"https://github.com/weihaox/VINDEX/blob/HEAD/src/model/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"cddbdad94b8e10ad"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}