Datasets › MediBeng
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset
The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings. The dataset is available under the CC-BY-4.0 license.
- Access the dataset on Hugging Face.
- For detailed instructions on dataset creation and storage, visit the GitHub Repository
Dataset Details:
- Number of Audio Files: 4800
- Total Duration: 7.11 hours
- Number of Speakers: 2 (1 Male, 1 Female)
- Utterance Pitch Mean: 335 - 673 Hz
- Utterance Pitch Standard Deviation: 210 - 493 Hz
- Sampling Rate: 16000 Hz
- Data Split: Train and Test
- Duration Range: 3.71s - 6.98s
- Languages: Code-mixed Bengali-English
- Gender Distribution: 1 Male, 1 Female
- Total File Size: 324 MB
- Speech Type: Medical-related
- Data Type: Synthetic
- Languages: Bengali, English
- Tasks: ASR, TTS, Machine Translation
- Context: Clinical (Healthcare)
- License: CC-BY-4.0
MediBeng Dataset Columns
- audio: Synthetic Bengali-English clinical conversations.
- text: Code-switched Bengali-English conversations.
- translation: English translation.
- speaker_name: Speaker's gender (e.g., Male, Female).
- utterance_pitch_mean: Mean pitch of the audio in Hertz (Hz).
- utterance_pitch_std: Pitch variation (standard deviation in Hertz).
Dataset Creation:
- Audio Collection: Conversations in Bengali-English for healthcare.
- Transcription: Code-switched sentences.
- Translation: Code-switched sentences English translation.
- Feature Engineering: Calculating pitch features.
- Storage: Available in Parquet format on Hugging Face.
Citation:
bibtex
@misc{promila_ghosh_2025,
author = {Promila Ghosh},
title = {MediBeng (Revision b05b594)},
year = 2025,
url = {https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng},
doi = {10.57967/hf/5187},
publisher = {Hugging Face}
}
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Speech-to-Text Translation | MediBeng | MediBeng Whisper Tiny Bleu 0.98 | MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED... | pr0mila/MediBeng-Whisper-Tiny | 2 | Compare |
Papers archive 2025-07-28
1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED BENGALI-ENGLISH TRANSLATOR FOR CLINICAL APPLICATIONS | 1 | 2 | 25 Apr 2025 | not harvested |
Dataset loaders archive 2025-07-28
2 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- MediBeng
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections