{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lhgnn-local-higher-order-graph-neural","title":"LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging","arxiv_id":"2501.03464","date":"2025-01-07","proceeding":null,"authors":["Shubhr Singh","Emmanouil Benetos","Huy Phan","Dan Stowell"],"abstract":"Transformers have set new benchmarks in audio processing tasks, leveraging self-attention mechanisms to capture complex patterns and dependencies within audio data. However, their focus on pairwise interactions limits their ability to process the higher-order relations essential for identifying distinct audio objects. To address this limitation, this work introduces the Local- Higher Order Graph Neural Network (LHGNN), a graph based model that enhances feature understanding by integrating local neighbourhood information with higher-order data from Fuzzy C-Means clusters, thereby capturing a broader spectrum of audio relationships. Evaluation of the model on three publicly available audio datasets shows that it outperforms Transformer-based models across all benchmarks while operating with substantially fewer parameters. Moreover, LHGNN demonstrates a distinct advantage in scenarios lacking ImageNet pretraining, establishing its effectiveness and efficiency in environments where extensive pretraining data is unavailable.","url_abs":"https://arxiv.org/abs/2501.03464v2","url_pdf":"https://arxiv.org/pdf/2501.03464v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"audio-classification","task_name":"Audio Classification"},{"task_slug":"graph-neural-network","task_name":"Graph Neural Network"}],"methods":[{"method_slug":"focus","method_name":"Focus"},{"method_slug":"graph-neural-network","method_name":"Graph Neural Network"},{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-classification-on-audio-set","task":"Audio Classification","dataset":"Audio Set","model":"LHGNN","rank_in_archive_order":2,"of":3,"metrics":{"Mean AP":"46.6"},"uses_additional_data":false},{"leaderboard":"/sota/audio-classification-on-esc-50","task":"Audio Classification","dataset":"ESC-50","model":"LHGNN","rank_in_archive_order":12,"of":29,"metrics":{"Top-1 Accuracy":"96.2"},"uses_additional_data":false},{"leaderboard":"/sota/audio-classification-on-fsd50k","task":"Audio Classification","dataset":"FSD50K","model":"LHGNN","rank_in_archive_order":10,"of":10,"metrics":{"Mean AP":"59"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}