{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-resource-light-method-for-cross-lingual","title":"A Resource-Light Method for Cross-Lingual Semantic Textual Similarity","arxiv_id":"1801.06436","date":"2018-01-19","proceeding":null,"authors":["Goran Glavaš","Marc Franco-Salvador","Simone Paolo Ponzetto","Paolo Rosso"],"abstract":"Recognizing semantically similar sentences or paragraphs across languages is\nbeneficial for many tasks, ranging from cross-lingual information retrieval and\nplagiarism detection to machine translation. Recently proposed methods for\npredicting cross-lingual semantic similarity of short texts, however, make use\nof tools and resources (e.g., machine translation systems, syntactic parsers or\nnamed entity recognition) that for many languages (or language pairs) do not\nexist. In contrast, we propose an unsupervised and a very resource-light\napproach for measuring semantic similarity between texts in different\nlanguages. To operate in the bilingual (or multilingual) space, we project\ncontinuous word vectors (i.e., word embeddings) from one language to the vector\nspace of the other language via the linear translation model. We then align\nwords according to the similarity of their vectors in the bilingual embedding\nspace and investigate different unsupervised measures of semantic similarity\nexploiting bilingual embeddings and word alignments. Requiring only a\nlimited-size set of word translation pairs between the languages, the proposed\napproach is applicable to virtually any pair of languages for which there\nexists a sufficiently large corpus, required to learn monolingual word\nembeddings. Experimental results on three different datasets for measuring\nsemantic textual similarity show that our simple resource-light approach\nreaches performance close to that of supervised and resource intensive methods,\ndisplaying stability across different language pairs. Furthermore, we evaluate\nthe proposed method on two extrinsic tasks, namely extraction of parallel\nsentences from comparable corpora and cross lingual plagiarism detection, and\nshow that it yields performance comparable to those of complex\nresource-intensive state-of-the-art models for the respective tasks.","url_abs":"http://arxiv.org/abs/1801.06436v1","url_pdf":"http://arxiv.org/pdf/1801.06436v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-resource-light-method-for-cross-lingual","repo_url":"https://bitbucket.org/gg42554/cl-sts","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"cross-lingual-information-retrieval","task_name":"Cross-Lingual Information Retrieval"},{"task_slug":"cross-lingual-semantic-textual-similarity","task_name":"Cross-Lingual Semantic Textual Similarity"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"word-translation","task_name":"Word Translation"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}