{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/character-level-neural-translation-for","title":"Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project","arxiv_id":"1604.01221","date":"2016-04-05","proceeding":"LREC 2016 5","authors":["Guntis Barzdins","Steve Renals","Didzis Gosko"],"abstract":"The paper steps outside the comfort-zone of the traditional NLP tasks like\nautomatic speech recognition (ASR) and machine translation (MT) to addresses\ntwo novel problems arising in the automated multilingual news monitoring:\nsegmentation of the TV and radio program ASR transcripts into individual\nstories, and clustering of the individual stories coming from various sources\nand languages into storylines. Storyline clustering of stories covering the\nsame events is an essential task for inquisitorial media monitoring. We address\nthese two problems jointly by engaging the low-dimensional semantic\nrepresentation capabilities of the sequence to sequence neural translation\nmodels. To enable joint multi-task learning for multilingual neural translation\nof morphologically rich languages we replace the attention mechanism with the\nsliding-window mechanism and operate the sequence to sequence neural\ntranslation model on the character-level rather than on the word-level. The\nstory segmentation and storyline clustering problem is tackled by examining the\nlow-dimensional vectors produced as a side-product of the neural translation\nprocess. The results of this paper describe a novel approach to the automatic\nstory segmentation and storyline clustering problem.","url_abs":"http://arxiv.org/abs/1604.01221v1","url_pdf":"http://arxiv.org/pdf/1604.01221v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"character-level-neural-translation-for","repo_url":"https://github.com/didzis/tensorflowAMR","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}