{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multilingual-distributed-representations","title":"Multilingual Distributed Representations without Word Alignment","arxiv_id":"1312.6173","date":"2013-12-20","proceeding":null,"authors":["Karl Moritz Hermann","Phil Blunsom"],"abstract":"Distributed representations of meaning are a natural way to encode covariance\nrelationships between words and phrases in NLP. By overcoming data sparsity\nproblems, as well as providing information about semantic relatedness which is\nnot available in discrete representations, distributed representations have\nproven useful in many NLP tasks. Recent work has shown how compositional\nsemantic representations can successfully be applied to a number of monolingual\napplications such as sentiment analysis. At the same time, there has been some\ninitial success in work on learning shared word-level representations across\nlanguages. We combine these two approaches by proposing a method for learning\ndistributed representations in a multilingual setup. Our model learns to assign\nsimilar embeddings to aligned sentences and dissimilar ones to sentence which\nare not aligned while not requiring word alignments. We show that our\nrepresentations are semantically informative and apply them to a cross-lingual\ndocument classification task where we outperform the previous state of the art.\nFurther, by employing parallel corpora of multiple language pairs we find that\nour model learns representations that capture semantic relationships across\nlanguages for which no parallel data was used.","url_abs":"http://arxiv.org/abs/1312.6173v4","url_pdf":"http://arxiv.org/pdf/1312.6173v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multilingual-distributed-representations","repo_url":"https://github.com/karlmoritz/bicvm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"cross-lingual-document-classification","task_name":"Cross-Lingual Document Classification"},{"task_slug":"document-classification","task_name":"Document Classification"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"},{"task_slug":"word-alignment","task_name":"Word Alignment"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-lingual-document-classification-on-12","task":"Cross-Lingual Document Classification","dataset":"Reuters RCV1/RCV2 English-to-German","model":"biCVM+","rank_in_archive_order":3,"of":3,"metrics":{"Accuracy":"86.2"},"uses_additional_data":false},{"leaderboard":"/sota/cross-lingual-document-classification-on-13","task":"Cross-Lingual Document Classification","dataset":"Reuters RCV1/RCV2 German-to-English","model":"biCVM+","rank_in_archive_order":3,"of":3,"metrics":{"Accuracy":"76.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1312.6173","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}