{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/speaker-recognition-from-raw-waveform-with","title":"Speaker Recognition from Raw Waveform with SincNet","arxiv_id":"1808.00158","date":"2018-07-29","proceeding":null,"authors":["Mirco Ravanelli","Yoshua Bengio"],"abstract":"Deep learning is progressively gaining popularity as a viable alternative to i-vectors for speaker recognition. Promising results have been recently obtained with Convolutional Neural Networks (CNNs) when fed by raw speech samples directly. Rather than employing standard hand-crafted features, the latter CNNs learn low-level speech representations from waveforms, potentially allowing the network to better capture important narrow-band speaker characteristics such as pitch and formants. Proper design of the neural network is crucial to achieve this goal. This paper proposes a novel CNN architecture, called SincNet, that encourages the first convolutional layer to discover more meaningful filters. SincNet is based on parametrized sinc functions, which implement band-pass filters. In contrast to standard CNNs, that learn all elements of each filter, only low and high cutoff frequencies are directly learned from data with the proposed method. This offers a very compact and efficient way to derive a customized filter bank specifically tuned for the desired application. Our experiments, conducted on both speaker identification and speaker verification tasks, show that the proposed architecture converges faster and performs better than a standard CNN on raw waveforms.","url_abs":"https://arxiv.org/abs/1808.00158v3","url_pdf":"https://arxiv.org/pdf/1808.00158v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/mravanelli/SincNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/008karan/SincNet_demo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/AI-Guru/SincNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/AryaAftab/sincnet-tensorflow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/HHousen/speaker-change-detection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/MarvinLvn/voice-type-classifier","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/ShristiShrestha/SincConvBasedSpeakerRecognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/Tonmoy1321/SincNet-499-Research","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/V-ivek/keras-sincnet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/busytex/busyide","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/drova326/sincnet_metric_learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/fbravosanchez/NIPS4Bplus","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/google-research/leaf-audio","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/grausof/keras-sincnet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/icewing1996/SincNet-for-Earthquake","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/icewing1996/SincNet_MLP","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/icewing1996/SincNet_asr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/jaehwlee/MuSincNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/jennyqsun/EEG-Decision-SincNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/juanmc2005/SpeakerEmbeddingLossComparison","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/kacper1095/speaker-gender-classification-and-deepdream","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/pkwww/SincNet_MLP","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/sajabdoli/UAP","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/shangeth/wavencoder","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/tuananh0305/End2End_SpeakRecognition_SincNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"speaker-recognition-from-raw-waveform-with","repo_url":"https://github.com/zzmtsvv/snzs","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"speaker-identification","task_name":"Speaker Identification"},{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"},{"task_slug":"speaker-verification","task_name":"Speaker Verification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.00158","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1808.00158"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Tonmoy1321/SincNet-499-Research","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jennyqsun/EEG-Decision-SincNet","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pkwww/SincNet_MLP","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/icewing1996/SincNet_asr","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jaehwlee/MuSincNet","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/AI-Guru/SincNet","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/HHousen/speaker-change-detection","reach":{"status":"ok","spdx":"GPL-3.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tuananh0305/End2End_SpeakRecognition_SincNet","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/drova326/sincnet_metric_learning","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/AryaAftab/sincnet-tensorflow","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/icewing1996/SincNet_MLP","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sajabdoli/UAP","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/google-research/leaf-audio","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ShristiShrestha/SincConvBasedSpeakerRecognition","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/fbravosanchez/NIPS4Bplus","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/shangeth/wavencoder","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/icewing1996/SincNet-for-Earthquake","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/008karan/SincNet_demo","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/busytex/busyide","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/juanmc2005/SpeakerEmbeddingLossComparison","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mravanelli/SincNet","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zzmtsvv/snzs","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/grausof/keras-sincnet","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MarvinLvn/voice-type-classifier","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/V-ivek/keras-sincnet","reach":{"status":"ok"}}],"summary":{"ran_draft_wrong":6,"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1},"listed":{"samples":6,"ran":6,"repositories":4}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"820ca1098b7fe1b7","entry":"act_fun","repo":"pkwww/SincNet_MLP","repo_kind":"listed","path":"dnn_models.py","file_url":"https://github.com/pkwww/SincNet_MLP/blob/HEAD/dnn_models.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"820ca1098b7fe1b7"}},{"code_sha256_prefix":"25b50b526d156bc2","entry":"act_fun","repo":"drova326/sincnet_metric_learning","repo_kind":"listed","path":"dnn_models.py","file_url":"https://github.com/drova326/sincnet_metric_learning/blob/HEAD/dnn_models.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"25b50b526d156bc2"}},{"code_sha256_prefix":"b380f42b5b71cc47","entry":"flip","repo":"pkwww/SincNet_MLP","repo_kind":"listed","path":"dnn_models.py","file_url":"https://github.com/pkwww/SincNet_MLP/blob/HEAD/dnn_models.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b380f42b5b71cc47"}},{"code_sha256_prefix":"43b56c6fdb708d63","entry":"ig_f","repo":"tuananh0305/End2End_SpeakRecognition_SincNet","repo_kind":"listed","path":"TIMIT_preparation.py","file_url":"https://github.com/tuananh0305/End2End_SpeakRecognition_SincNet/blob/HEAD/TIMIT_preparation.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"43b56c6fdb708d63"}},{"code_sha256_prefix":"a082e26af71d25c2","entry":"ig_f","repo":"zzmtsvv/snzs","repo_kind":"listed","path":"Sincnet.py","file_url":"https://github.com/zzmtsvv/snzs/blob/HEAD/Sincnet.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a082e26af71d25c2"}},{"code_sha256_prefix":"b31437bae9320e91","entry":"sinc","repo":"pkwww/SincNet_MLP","repo_kind":"listed","path":"dnn_models.py","file_url":"https://github.com/pkwww/SincNet_MLP/blob/HEAD/dnn_models.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b31437bae9320e91"}},{"code_sha256_prefix":"bf4e6d63364092e1","entry":"ReadList","repo":"mravanelli/SincNet","repo_kind":"official","path":"data_io.py","file_url":"https://github.com/mravanelli/SincNet/blob/HEAD/data_io.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bf4e6d63364092e1"}},{"code_sha256_prefix":"92a2612cf623f10e","entry":"create_batches_rnd","repo":"mravanelli/SincNet","repo_kind":"official","path":"data_io.py","file_url":"https://github.com/mravanelli/SincNet/blob/HEAD/data_io.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"92a2612cf623f10e"}},{"code_sha256_prefix":"fb6bced9440ef6c6","entry":"str_to_bool","repo":"mravanelli/SincNet","repo_kind":"official","path":"data_io.py","file_url":"https://github.com/mravanelli/SincNet/blob/HEAD/data_io.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fb6bced9440ef6c6"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}