{"url":"/dataset/callhome-american-english-speech","name":"CALLHOME American English Speech","full_name":null,"description_markdown":"The CALLHOME English Corpus is a collection of unscripted telephone conversations between native speakers of English. Here are the key details:\r\n\r\nParticipants: 120 individuals.\r\nType of Study: Naturalistic.\r\nLocation: USA.\r\nMedia Type: Audio.\r\nDOI: doi:10.21415/T5KP54.\r\nContents: The corpus contains 120 telephone conversations, each lasting up to 30 minutes.\r\nOrigins: All calls originated in North America, with 90 calls placed to various locations overseas and 30 calls within North America.\r\nTranscripts: The transcripts cover contiguous 5 or 10-minute segments from recorded conversations.\r\nSpeaker Awareness: All speakers were aware that they were being recorded.\r\nTopics: Participants had no guidelines on what to talk about; most called family members or close friends overseas.\r\nPurpose: Collected primarily to support the project on Large Vocabulary Conversational Speech Recognition (LVCSR), sponsored by the U.S. Department of Defense.","description_withheld":null,"homepage":"https://catalog.ldc.upenn.edu/LDC97S42","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[{"name":"Speech Recognition","url":"/task/speech-recognition","datasets_with_task":"/datasets/task/speech-recognition"},{"name":"Automatic Speech Recognition","url":"/task/automatic-speech-recognition-2","datasets_with_task":"/datasets/task/automatic-speech-recognition-2"},{"name":"Speaker Diarization","url":"/task/speaker-diarization","datasets_with_task":"/datasets/task/speaker-diarization"},{"name":"Speaker Verification","url":"/task/speaker-verification","datasets_with_task":"/datasets/task/speaker-verification"}],"languages":[],"variants":["Switchboard CallHome","Hub5'00 CallHome","CALLHOME Spanish Speech (Test)","CALLHOME Spanish Speech (Dev)","CALLHOME Spanish Speech","CALLHOME-109","CALLHOME","CALLHOME American English Speech"],"data_loaders":[],"num_papers_in_archive":11,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/speaker-diarization-on-callhome","task":"Speaker Diarization","dataset_variant":"CALLHOME","rows":10,"metrics":["DER(%)","DER(ig olp)","FA","MI","CF"],"first_row_in_archive_order":{"model":"TOLD","paper":"/paper/told-a-novel-two-stage-overlap-aware","metrics":{"CF":"2.94","DER(%)":"10.14","DER(ig olp)":"7.37","FA":"2.4","MI":"4.8"},"code_links":[{"title":"alibaba-damo-academy/FunASR","url":"https://github.com/alibaba-damo-academy/FunASR"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speaker-diarization-on-callhome-109","task":"Speaker Diarization","dataset_variant":"CALLHOME-109","rows":2,"metrics":["DER(%)"],"first_row_in_archive_order":{"model":"titanet-s","paper":"/paper/titanet-neural-model-for-speaker","metrics":{"DER(%)":"1.11"},"code_links":[{"title":"NVIDIA/NeMo","url":"https://github.com/NVIDIA/NeMo"},{"title":"Wadaboa/titanet","url":"https://github.com/Wadaboa/titanet"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speaker-verification-on-callhome","task":"Speaker Verification","dataset_variant":"CALLHOME","rows":2,"metrics":["Cosine EER"],"first_row_in_archive_order":{"model":"GE2E","paper":"/paper/generalized-end-to-end-loss-for-speaker","metrics":{"Cosine EER":"3.55"},"code_links":[{"title":"CorentinJ/Real-Time-Voice-Cloning","url":"https://github.com/CorentinJ/Real-Time-Voice-Cloning"},{"title":"coqui-ai/TTS","url":"https://github.com/coqui-ai/TTS"},{"title":"PaddlePaddle/PaddleSpeech","url":"https://github.com/PaddlePaddle/PaddleSpeech"},{"title":"resemble-ai/Resemblyzer","url":"https://github.com/resemble-ai/Resemblyzer"},{"title":"HarryVolek/PyTorch_Speaker_Verification","url":"https://github.com/HarryVolek/PyTorch_Speaker_Verification"},{"title":"google/speaker-id","url":"https://github.com/google/speaker-id/tree/master/lingvo"},{"title":"JanhHyun/Speaker_Verification","url":"https://github.com/JanhHyun/Speaker_Verification"},{"title":"Janghyun1230/Speaker_Verification","url":"https://github.com/Janghyun1230/Speaker_Verification"},{"title":"yistLin/dvector","url":"https://github.com/yistLin/dvector"},{"title":"rf5/simple-speaker-embedding","url":"https://github.com/rf5/simple-speaker-embedding"},{"title":"cvqluu/GE2E-Loss","url":"https://github.com/cvqluu/GE2E-Loss"},{"title":"Suhee05/Text-Independent-Speaker-Verification","url":"https://github.com/Suhee05/Text-Independent-Speaker-Verification"},{"title":"Aurora11111/voiceprint","url":"https://github.com/Aurora11111/voiceprint"},{"title":"Aurora11111/speaker-recognition-pytorch","url":"https://github.com/Aurora11111/speaker-recognition-pytorch"},{"title":"muskang48/Speaker-Diarization","url":"https://github.com/muskang48/Speaker-Diarization"},{"title":"tigthor/Voice-Cloning-AI","url":"https://github.com/tigthor/Voice-Cloning-AI"},{"title":"piotrkawa/audio-deepfake-source-tracing","url":"https://github.com/piotrkawa/audio-deepfake-source-tracing"},{"title":"JeffT13/rd-diarization","url":"https://github.com/JeffT13/rd-diarization"},{"title":"gkv856/speaker_embedding_GE2E_loss","url":"https://github.com/gkv856/speaker_embedding_GE2E_loss"},{"title":"dalonlobo/diarization-experiments","url":"https://github.com/dalonlobo/diarization-experiments"},{"title":"luomingshuang/GE2E-SV-TI-Voxceleb-LMS","url":"https://github.com/luomingshuang/GE2E-SV-TI-Voxceleb-LMS"},{"title":"yui-mhcp/base_dl_project","url":"https://github.com/yui-mhcp/base_dl_project"},{"title":"luomingshuang/GE2E-SV-TI-Timit-LMS","url":"https://github.com/luomingshuang/GE2E-SV-TI-Timit-LMS"},{"title":"aijianiula0601/ge2eloss-svf","url":"https://github.com/aijianiula0601/ge2eloss-svf"},{"title":"luomingshuang/GE2E-SV-TI-thchs30-LMS","url":"https://github.com/luomingshuang/GE2E-SV-TI-thchs30-LMS"},{"title":"JeffT13/VoiceEncoder","url":"https://github.com/JeffT13/VoiceEncoder"},{"title":"hanqingguo/GE2E","url":"https://github.com/hanqingguo/GE2E"},{"title":"zhangmin4215/PyTorch_Speaker_Verification","url":"https://github.com/zhangmin4215/PyTorch_Speaker_Verification"},{"title":"icewing1996/baseline","url":"https://github.com/icewing1996/baseline"},{"title":"luomingshuang/GE2E-SV-TI-Chinese-LMS","url":"https://github.com/luomingshuang/GE2E-SV-TI-Chinese-LMS"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speaker-diarization-on-hub5-00-callhome","task":"Speaker Diarization","dataset_variant":"Hub5'00 CallHome","rows":1,"metrics":["V"],"first_row_in_archive_order":{"model":"UIS-RNN","paper":"/paper/fully-supervised-speaker-diarization","metrics":{"V":"10.6"},"code_links":[{"title":"google/uis-rnn","url":"https://github.com/google/uis-rnn"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-recognition-on-callhome-spanish-speech","task":"Speech Recognition","dataset_variant":"CALLHOME Spanish Speech","rows":1,"metrics":["WER"],"first_row_in_archive_order":{"model":"TDT 0-2","paper":"/paper/efficient-sequence-transduction-by-jointly","metrics":{"WER":"17.95"},"code_links":[{"title":"NVIDIA/NeMo","url":"https://github.com/NVIDIA/NeMo"},{"title":"chimechallenge/C8DASR-Baseline-NeMo","url":"https://github.com/chimechallenge/C8DASR-Baseline-NeMo"},{"title":"kehanlu/Nemo","url":"https://github.com/kehanlu/Nemo"},{"title":"wd929/NeMo","url":"https://github.com/wd929/NeMo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-recognition-on-hub500-callhome","task":"Speech Recognition","dataset_variant":"Hub5'00 CallHome","rows":1,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"Espresso","paper":"/paper/espresso-a-fast-end-to-end-neural-speech","metrics":{"Word Error Rate (WER)":"19.1"},"code_links":[{"title":"freewym/espresso","url":"https://github.com/freewym/espresso"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-recognition-on-switchboard-callhome","task":"Speech Recognition","dataset_variant":"Switchboard CallHome","rows":1,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"SpeechStew (100M)","paper":"/paper/speechstew-simply-mix-all-available-speech","metrics":{"Word Error Rate (WER)":"8.3"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/efficient-sequence-transduction-by-jointly","title":"Efficient Sequence Transduction by Jointly Predicting Tokens and Durations","date":"2023-04-13","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":1,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/told-a-novel-two-stage-overlap-aware","title":"TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization","date":"2023-03-08","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/titanet-neural-model-for-speaker","title":"TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context","date":"2021-10-08","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":0,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/speechstew-simply-mix-all-available-speech","title":"SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network","date":"2021-04-05","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/auto-tuning-spectral-clustering-for-speaker","title":"Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap","date":"2020-03-05","rows_on_this_dataset":5,"code_links":1,"syntology":null},{"paper":"/paper/espresso-a-fast-end-to-end-neural-speech","title":"Espresso: A Fast End-to-end Neural Speech Recognition Toolkit","date":"2019-09-18","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/end-to-end-neural-speaker-diarization-with","title":"End-to-End Neural Speaker Diarization with Self-attention","date":"2019-09-13","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":0,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/end-to-end-neural-speaker-diarization-with-1","title":"End-to-End Neural Speaker Diarization with Permutation-Free Objectives","date":"2019-09-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/fully-supervised-speaker-diarization","title":"Fully Supervised Speaker Diarization","date":"2018-10-10","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/speaker-diarization-with-lstm","title":"Speaker Diarization with LSTM","date":"2017-10-28","rows_on_this_dataset":1,"code_links":4,"syntology":null},{"paper":"/paper/generalized-end-to-end-loss-for-speaker","title":"Generalized End-to-End Loss for Speaker Verification","date":"2017-10-28","rows_on_this_dataset":2,"code_links":30,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":27,"samples_ran":4,"samples_unverified":23,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":5,"samples_harvested":39,"samples_ran":6,"samples_unverified":33,"pointer_only_for_licence":2,"papers_with_no_sample_that_ran":2,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}