{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmcr4nlp-multilingual-multiway-corpora","title":"MMCR4NLP: Multilingual Multiway Corpora Repository for Natural Language Processing","arxiv_id":"1710.01025","date":"2017-10-03","proceeding":null,"authors":["Raj Dabre","Sadao Kurohashi"],"abstract":"Multilinguality is gradually becoming ubiquitous in the sense that more and\nmore researchers have successfully shown that using additional languages help\nimprove the results in many Natural Language Processing tasks. Multilingual\nMultiway Corpora (MMC) contain the same sentence in multiple languages. Such\ncorpora have been primarily used for Multi-Source and Pivot Language Machine\nTranslation but are also useful for developing multilingual sequence taggers by\ntransfer learning. While these corpora are available, they are not organized\nfor multilingual experiments and researchers need to write boilerplate code\nevery time they want to use said corpora. Moreover, because there is no\nofficial MMC collection it becomes difficult to compare against existing\napproaches. As such we present our work on creating a unified and\nsystematically organized repository of MMC spanning a large number of\nlanguages. We also provide training, development and test splits for corpora\nwhere official splits are unavailable. We hope that this will help speed up the\npace of multilingual NLP research and ensure that NLP researchers obtain\nresults that are more trustable since they can be compared easily. We indicate\ncorpora sources, extraction procedures if any and relevant statistics. We also\nmake our collection public for research purposes.","url_abs":"http://arxiv.org/abs/1710.01025v3","url_pdf":"http://arxiv.org/pdf/1710.01025v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmcr4nlp-multilingual-multiway-corpora","repo_url":"https://github.com/alphadl/languageaware_tuning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"mmcr4nlp-multilingual-multiway-corpora","repo_url":"https://github.com/zhiqu22/adapnoncenter","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"multilingual-nlp","task_name":"Multilingual NLP"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"translation","task_name":"Translation"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1710.01025","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}