{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/aenet-learning-deep-audio-features-for-video","title":"AENet: Learning Deep Audio Features for Video Analysis","arxiv_id":"1701.00599","date":"2017-01-03","proceeding":null,"authors":["Naoya Takahashi","Michael Gygli","Luc van Gool"],"abstract":"We propose a new deep network for audio event recognition, called AENet. In\ncontrast to speech, sounds coming from audio events may be produced by a wide\nvariety of sources. Furthermore, distinguishing them often requires analyzing\nan extended time period due to the lack of clear sub-word units that are\npresent in speech. In order to incorporate this long-time frequency structure\nof audio events, we introduce a convolutional neural network (CNN) operating on\na large temporal input. In contrast to previous works this allows us to train\nan audio event detection system end-to-end. The combination of our network\narchitecture and a novel data augmentation outperforms previous methods for\naudio event detection by 16%. Furthermore, we perform transfer learning and\nshow that our model learnt generic audio features, similar to the way CNNs\nlearn generic features on vision tasks. In video analysis, combining visual\nfeatures and traditional audio features such as MFCC typically only leads to\nmarginal improvements. Instead, combining visual features with our AENet\nfeatures, which can be computed efficiently on a GPU, leads to significant\nperformance improvements on action recognition and video highlight detection.\nIn video highlight detection, our audio features improve the performance by\nmore than 8% over visual features alone.","url_abs":"http://arxiv.org/abs/1701.00599v2","url_pdf":"http://arxiv.org/pdf/1701.00599v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"aenet-learning-deep-audio-features-for-video","repo_url":"https://github.com/znaoya/aenet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"event-detection","task_name":"Event Detection"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"highlight-detection","task_name":"Highlight Detection"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}