Datasets › MediBeng

MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)

14 Apr 2025 archive 2025-07-28

MediBeng Dataset

The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings. The dataset is available under the CC-BY-4.0 license.

Dataset Details:
  • Number of Audio Files: 4800
  • Total Duration: 7.11 hours
  • Number of Speakers: 2 (1 Male, 1 Female)
  • Utterance Pitch Mean: 335 - 673 Hz
  • Utterance Pitch Standard Deviation: 210 - 493 Hz
  • Sampling Rate: 16000 Hz
  • Data Split: Train and Test
  • Duration Range: 3.71s - 6.98s
  • Languages: Code-mixed Bengali-English
  • Gender Distribution: 1 Male, 1 Female
  • Total File Size: 324 MB
  • Speech Type: Medical-related
  • Data Type: Synthetic
  • Languages: Bengali, English
  • Tasks: ASR, TTS, Machine Translation
  • Context: Clinical (Healthcare)
  • License: CC-BY-4.0
MediBeng Dataset Columns
  • audio: Synthetic Bengali-English clinical conversations.
  • text: Code-switched Bengali-English conversations.
  • translation: English translation.
  • speaker_name: Speaker's gender (e.g., Male, Female).
  • utterance_pitch_mean: Mean pitch of the audio in Hertz (Hz).
  • utterance_pitch_std: Pitch variation (standard deviation in Hertz).
Dataset Creation:
  1. Audio Collection: Conversations in Bengali-English for healthcare.
  2. Transcription: Code-switched sentences.
  3. Translation: Code-switched sentences English translation.
  4. Feature Engineering: Calculating pitch features.
  5. Storage: Available in Parquet format on Hugging Face.
Citation:
bibtex
@misc{promila_ghosh_2025,
    author       = {Promila Ghosh},
    title        = {MediBeng (Revision b05b594)},
    year         = 2025,
    url          = {https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng},
    doi          = {10.57967/hf/5187},
    publisher    = {Hugging Face}
}

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speech-to-Text Translation MediBeng MediBeng Whisper Tiny Bleu 0.98 MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED... pr0mila/MediBeng-Whisper-Tiny 2 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED BENGALI-ENGLISH TRANSLATOR FOR CLINICAL APPLICATIONS 1 2 25 Apr 2025 not harvested

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC-BY-4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MediBeng

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections