Datasets › CALLHOME American English Speech

CALLHOME American English Speech

archive 2025-07-28

The CALLHOME English Corpus is a collection of unscripted telephone conversations between native speakers of English. Here are the key details:

Participants: 120 individuals. Type of Study: Naturalistic. Location: USA. Media Type: Audio. DOI: doi:10.21415/T5KP54. Contents: The corpus contains 120 telephone conversations, each lasting up to 30 minutes. Origins: All calls originated in North America, with 90 calls placed to various locations overseas and 30 calls within North America. Transcripts: The transcripts cover contiguous 5 or 10-minute segments from recorded conversations. Speaker Awareness: All speakers were aware that they were being recorded. Topics: Participants had no guidelines on what to talk about; most called family members or close friends overseas. Purpose: Collected primarily to support the project on Large Vocabulary Conversational Speech Recognition (LVCSR), sponsored by the U.S. Department of Defense.

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 11. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations 4 1 13 Apr 2023 ran 1 of 2 samples (1 unverified)
TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization 1 2 8 Mar 2023 not harvested
TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context 2 1 8 Oct 2021 ran 0 of 4 samples (4 unverified)
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network 0 1 5 Apr 2021 not harvested
Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap 1 5 5 Mar 2020 not harvested
Espresso: A Fast End-to-end Neural Speech Recognition Toolkit 1 1 18 Sep 2019 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
End-to-End Neural Speaker Diarization with Self-attention 2 2 13 Sep 2019 ran 0 of 5 samples (5 unverified)
End-to-End Neural Speaker Diarization with Permutation-Free Objectives 1 1 12 Sep 2019 not harvested
Fully Supervised Speaker Diarization 1 1 10 Oct 2018 not harvested
Speaker Diarization with LSTM 4 1 28 Oct 2017 not harvested
Generalized End-to-End Loss for Speaker Verification 30 2 28 Oct 2017 ran 4 of 27 samples (23 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Switchboard CallHome
  • Hub5'00 CallHome
  • CALLHOME Spanish Speech (Test)
  • CALLHOME Spanish Speech (Dev)
  • CALLHOME Spanish Speech
  • CALLHOME-109
  • CALLHOME
  • CALLHOME American English Speech

8 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections