{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/very-deep-convolutional-neural-networks-for","title":"Very Deep Convolutional Neural Networks for Raw Waveforms","arxiv_id":"1610.00087","date":"2016-10-01","proceeding":null,"authors":["Wei Dai","Chia Dai","Shuhui Qu","Juncheng Li","Samarjit Das"],"abstract":"Learning acoustic models directly from the raw waveform data with minimal\nprocessing is challenging. Current waveform-based models have generally used\nvery few (~2) convolutional layers, which might be insufficient for building\nhigh-level discriminative features. In this work, we propose very deep\nconvolutional neural networks (CNNs) that directly use time-domain waveforms as\ninputs. Our CNNs, with up to 34 weight layers, are efficient to optimize over\nvery long sequences (e.g., vector of size 32000), necessary for processing\nacoustic waveforms. This is achieved through batch normalization, residual\nlearning, and a careful design of down-sampling in the initial layers. Our\nnetworks are fully convolutional, without the use of fully connected layers and\ndropout, to maximize representation learning. We use a large receptive field in\nthe first convolutional layer to mimic bandpass filters, but very small\nreceptive fields subsequently to control the model capacity. We demonstrate the\nperformance gains with the deeper models. Our evaluation shows that the CNN\nwith 18 weight layers outperform the CNN with 3 weight layers by over 15% in\nabsolute accuracy for an environmental sound recognition task and matches the\nperformance of models using log-mel features.","url_abs":"http://arxiv.org/abs/1610.00087v1","url_pdf":"http://arxiv.org/pdf/1610.00087v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/Andry-Heritiana/Very_Deep_CNN_Audio_Classification","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/am-sirdaniel/Audio_WaveForm_Paper_Implementation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/danielajisafe/audio_waveform_paper_implementation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/djmMax/audio-classification","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":{"status":"ok"}},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/francisanokye/Audio-Signal-Processing","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/lkidane/Very-Deep-Convolutional-Networks-For-Raw-Waveforms-pytorch-implementation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/nathoe/Audio-Classifier","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/philipperemy/very-deep-convnets-raw-waveforms","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/uman95/AMMI_research","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"very-deep-convolutional-neural-networks-for","repo_url":"https://github.com/vivek081166/raw-audio-deep-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1610.00087","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1610.00087"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vivek081166/raw-audio-deep-learning","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/am-sirdaniel/Audio_WaveForm_Paper_Implementation","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lkidane/Very-Deep-Convolutional-Networks-For-Raw-Waveforms-pytorch-implementation","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/djmMax/audio-classification","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Andry-Heritiana/Very_Deep_CNN_Audio_Classification","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/danielajisafe/audio_waveform_paper_implementation","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/francisanokye/Audio-Signal-Processing","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/uman95/AMMI_research","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/philipperemy/very-deep-convnets-raw-waveforms","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/nathoe/Audio-Classifier","reach":{"status":"ok"}}],"summary":{"ran_honours":1,"ran_draft_wrong":1,"unverified":7},"by_repo_kind":{"listed":{"samples":9,"ran":2,"repositories":4}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"a845b9c00079e7c8","entry":"extract_class_id","repo":"vivek081166/raw-audio-deep-learning","repo_kind":"listed","path":"process_data.py","file_url":"https://github.com/vivek081166/raw-audio-deep-learning/blob/HEAD/process_data.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a845b9c00079e7c8"}},{"code_sha256_prefix":"98695a53bf929f5b","entry":"get_data","repo":"vivek081166/raw-audio-deep-learning","repo_kind":"listed","path":"model.py","file_url":"https://github.com/vivek081166/raw-audio-deep-learning/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"98695a53bf929f5b"}},{"code_sha256_prefix":"688c38ac8517471a","entry":"accuracy","repo":"francisanokye/Audio-Signal-Processing","repo_kind":"listed","path":"Signal_Processing_of_Waveforms.py","file_url":"https://github.com/francisanokye/Audio-Signal-Processing/blob/HEAD/Signal_Processing_of_Waveforms.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"688c38ac8517471a"}},{"code_sha256_prefix":"3efc9b6319c29caa","entry":"extract_class_id","repo":"philipperemy/very-deep-convnets-raw-waveforms","repo_kind":"listed","path":"process_data.py","file_url":"https://github.com/philipperemy/very-deep-convnets-raw-waveforms/blob/HEAD/process_data.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3efc9b6319c29caa"}},{"code_sha256_prefix":"6d04ede5be0e3d1b","entry":"identity_block","repo":"philipperemy/very-deep-convnets-raw-waveforms","repo_kind":"listed","path":"model_resnet.py","file_url":"https://github.com/philipperemy/very-deep-convnets-raw-waveforms/blob/HEAD/model_resnet.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6d04ede5be0e3d1b"}},{"code_sha256_prefix":"c8612b8ec038c334","entry":"next_batch_blank","repo":"philipperemy/very-deep-convnets-raw-waveforms","repo_kind":"listed","path":"model_data.py","file_url":"https://github.com/philipperemy/very-deep-convnets-raw-waveforms/blob/HEAD/model_data.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c8612b8ec038c334"}},{"code_sha256_prefix":"e717a0964a471163","entry":"res_upsamp","repo":"Andry-Heritiana/Very_Deep_CNN_Audio_Classification","repo_kind":"listed","path":"audio_classifier.py","file_url":"https://github.com/Andry-Heritiana/Very_Deep_CNN_Audio_Classification/blob/HEAD/audio_classifier.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e717a0964a471163"}},{"code_sha256_prefix":"8073a6d3ba5a0498","entry":"test","repo":"Andry-Heritiana/Very_Deep_CNN_Audio_Classification","repo_kind":"listed","path":"audio_classifier.py","file_url":"https://github.com/Andry-Heritiana/Very_Deep_CNN_Audio_Classification/blob/HEAD/audio_classifier.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8073a6d3ba5a0498"}},{"code_sha256_prefix":"10bcc83b10a601ad","entry":"train","repo":"Andry-Heritiana/Very_Deep_CNN_Audio_Classification","repo_kind":"listed","path":"audio_classifier.py","file_url":"https://github.com/Andry-Heritiana/Very_Deep_CNN_Audio_Classification/blob/HEAD/audio_classifier.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"10bcc83b10a601ad"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}