{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spoken-english-intelligibility-remediation","title":"Spoken English Intelligibility Remediation with PocketSphinx Alignment and Feature Extraction Improves Substantially over the State of the Art","arxiv_id":"1709.01713","date":"2017-09-06","proceeding":null,"authors":["Yuan Gao","Brij Mohan Lal Srivastava","James Salsman"],"abstract":"We use automatic speech recognition to assess spoken English learner\npronunciation based on the authentic intelligibility of the learners' spoken\nresponses determined from support vector machine (SVM) classifier or deep\nlearning neural network model predictions of transcription correctness. Using\nnumeric features produced by PocketSphinx alignment mode and many recognition\npasses searching for the substitution and deletion of each expected phoneme and\ninsertion of unexpected phonemes in sequence, the SVM models achieve 82 percent\nagreement with the accuracy of Amazon Mechanical Turk crowdworker\ntranscriptions, up from 75 percent reported by multiple independent\nresearchers. Using such features with SVM classifier probability prediction\nmodels can help computer-aided pronunciation teaching (CAPT) systems provide\nintelligibility remediation.","url_abs":"http://arxiv.org/abs/1709.01713v3","url_pdf":"http://arxiv.org/pdf/1709.01713v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spoken-english-intelligibility-remediation","repo_url":"https://github.com/jsalsman/featex","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"svm","method_name":"SVM"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}