{"url":"/dataset/medibeng","name":"MediBeng","full_name":"Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications","description_markdown":"## MediBeng Dataset\r\n\r\nThe **MediBeng** dataset contains synthetic **code-switched dialogues** in **Bengali and English** for training models in **speech recognition (ASR)**, **text-to-speech (TTS)**, and **machine translation** in clinical settings.  The dataset is available under the **CC-BY-4.0** license. \r\n\r\n- Access the dataset on [Hugging Face](https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng).\r\n- For detailed instructions on dataset creation and storage, visit the [GitHub Repository](https://github.com/pr0mila/ParquetToHuggingFace)\r\n\r\n### Dataset Details:\r\n\r\n- **Number of Audio Files**: 4800\r\n- **Total Duration**: 7.11 hours\r\n- **Number of Speakers**: 2 (1 Male, 1 Female)\r\n- **Utterance Pitch Mean**: 335 - 673 Hz\r\n- **Utterance Pitch Standard Deviation**: 210 - 493 Hz\r\n- **Sampling Rate**: 16000 Hz\r\n- **Data Split**: Train and Test\r\n- **Duration Range**: 3.71s - 6.98s\r\n- **Languages**: Code-mixed Bengali-English\r\n- **Gender Distribution**: 1 Male, 1 Female\r\n- **Total File Size**: 324 MB\r\n- **Speech Type**: Medical-related\r\n- **Data Type**: Synthetic\r\n- **Languages**: Bengali, English\r\n- **Tasks**: ASR, TTS, Machine Translation\r\n- **Context**: Clinical (Healthcare)\r\n- **License**: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)\r\n\r\n### MediBeng Dataset Columns\r\n- **audio**: Synthetic Bengali-English clinical conversations.\r\n- **text**: Code-switched Bengali-English conversations.\r\n- **translation**: English translation.\r\n- **speaker_name**: Speaker's gender (e.g., Male, Female).\r\n- **utterance_pitch_mean**: Mean pitch of the audio in Hertz (Hz).\r\n- **utterance_pitch_std**: Pitch variation (standard deviation in Hertz).\r\n\r\n### Dataset Creation:\r\n1. **Audio Collection**: Conversations in Bengali-English for healthcare.\r\n2. **Transcription**: Code-switched sentences.\r\n3. **Translation**: Code-switched sentences English translation.\r\n4. **Feature Engineering**: Calculating pitch features.\r\n5. **Storage**: Available in **Parquet format** on [Hugging Face](https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng).\r\n\r\n### Citation:\r\n```\r\nbibtex\r\n@misc{promila_ghosh_2025,\r\n\tauthor       = {Promila Ghosh},\r\n\ttitle        = {MediBeng (Revision b05b594)},\r\n\tyear         = 2025,\r\n\turl          = {https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng},\r\n\tdoi          = {10.57967/hf/5187},\r\n\tpublisher    = {Hugging Face}\r\n}\r\n```","description_withheld":null,"homepage":"https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng","introduced_date":"2025-04-14","introduced_date_note":null,"introduced_by":null,"license":{"name":"CC-BY-4.0","url":"https://huggingface.co/datasets/choosealicense/licenses/blob/main/markdown/cc-by-4.0.md"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Audio","url":"/datasets/modality/audio"},{"name":"Medical","url":"/datasets/modality/medical"},{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Speech Recognition","url":"/task/speech-recognition","datasets_with_task":"/datasets/task/speech-recognition"},{"name":"Automatic Speech Recognition","url":"/task/automatic-speech-recognition-2","datasets_with_task":"/datasets/task/automatic-speech-recognition-2"},{"name":"Audio Classification","url":"/task/audio-classification","datasets_with_task":"/datasets/task/audio-classification"},{"name":"Speech-to-Text Translation","url":"/task/speech-to-text-translation","datasets_with_task":"/datasets/task/speech-to-text-translation"},{"name":"Audio Generation","url":"/task/audio-generation","datasets_with_task":"/datasets/task/audio-generation"},{"name":"Synthetic Data Generation","url":"/task/synthetic-data-generation","datasets_with_task":"/datasets/task/synthetic-data-generation"},{"name":"automatic-speech-translation","url":"/task/automatic-speech-translation","datasets_with_task":"/datasets/task/automatic-speech-translation"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"Bengali","url":"/datasets/language/bengali"}],"variants":["MediBeng"],"data_loaders":[{"repo":"https://github.com/pr0mila/ParquetToHuggingFace","url":"https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng","frameworks":[]},{"repo":"https://github.com/pr0mila/ParquetToHuggingFace","url":"https://github.com/pr0mila/ParquetToHuggingFace","frameworks":[]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/speech-to-text-translation-on-medibeng","task":"Speech-to-Text Translation","dataset_variant":"MediBeng","rows":2,"metrics":["Bleu"],"first_row_in_archive_order":{"model":"MediBeng Whisper Tiny","paper":"/paper/medibeng-whisper-tiny-a-fine-tuned-code","metrics":{"Bleu":"0.98"},"code_links":[{"title":"pr0mila/MediBeng-Whisper-Tiny","url":"https://github.com/pr0mila/MediBeng-Whisper-Tiny"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/medibeng-whisper-tiny-a-fine-tuned-code","title":"MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED BENGALI-ENGLISH TRANSLATOR FOR CLINICAL APPLICATIONS","date":"2025-04-25","rows_on_this_dataset":2,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}