Papers › MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED BENGALI-ENGLISH TRANSLATOR FOR...

MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED BENGALI-ENGLISH TRANSLATOR FOR CLINICAL APPLICATIONS

25 Apr 2025medRxiv 2025 4archive 2025-07-28

Promila Ghosh, Sunipun Talukder

Code-switching in multilingual healthcare settings challenges automated transcription systems as facilities adopt AI documentation tools. To tackle this issue, we developed a cost-effective solution using the MediBeng Whisper Tiny model, a fine-tuned version of the Whisper Tiny model specifically designed for code-switched Bengali-English conversations in healthcare contexts. The model was fine-tuned on MediBeng, a synthetic dataset created to simulate the types of bilingual interactions often found in healthcare environments. The fine-tuning process involved using just 20% of the dataset, making it highly efficient in terms of computational resources and data usage. Despite the limited data, our model achieved an exceptional 0.01 Word Error Rate (WER), reflecting near-perfect transcription accuracy. Additionally, the model attained a 0.98 BLEU score, indicating its ability to accurately translate mixed-language input into English. These impressive results demonstrate that even with minimal data and resources, the model can handle complex code-switching tasks effectively. This model improves information processing for healthcare professionals, allowing doctors to save time on paperwork and focus more on patient care. It also ensures patient records are more accurate and accessible, aiding better healthcare decision-making.

PaperPDFCode

Code

pr0mila/MediBeng-Whisper-Tiny mentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Clinical Language TranslationMachine TranslationSpeech-to-TextSpeech-to-Text TranslationSynthetic Data Generation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech-to-Text Translation MediBeng MediBeng Whisper Tiny Bleu 0.98 #1 of 2 Archive leaderboard report
Speech-to-Text Translation MediBeng Whisper Tiny Bleu 0.30 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ADOPTAbsolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutFocusLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections