{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-cross-lingual-information","title":"Unsupervised Cross-Lingual Information Retrieval using Monolingual Data Only","arxiv_id":"1805.00879","date":"2018-05-02","proceeding":null,"authors":["Robert Litschko","Goran Glavaš","Simone Paolo Ponzetto","Ivan Vulić"],"abstract":"We propose a fully unsupervised framework for ad-hoc cross-lingual\ninformation retrieval (CLIR) which requires no bilingual data at all. The\nframework leverages shared cross-lingual word embedding spaces in which terms,\nqueries, and documents can be represented, irrespective of their actual\nlanguage. The shared embedding spaces are induced solely on the basis of\nmonolingual corpora in two languages through an iterative process based on\nadversarial neural networks. Our experiments on the standard CLEF CLIR\ncollections for three language pairs of varying degrees of language similarity\n(English-Dutch/Italian/Finnish) demonstrate the usefulness of the proposed\nfully unsupervised approach. Our CLIR models with unsupervised cross-lingual\nembeddings outperform baselines that utilize cross-lingual embeddings induced\nrelying on word-level and document-level alignments. We then demonstrate that\nfurther improvements can be achieved by unsupervised ensemble CLIR models. We\nbelieve that the proposed framework is the first step towards development of\neffective CLIR models for language pairs and domains where parallel data are\nscarce or non-existent.","url_abs":"http://arxiv.org/abs/1805.00879v1","url_pdf":"http://arxiv.org/pdf/1805.00879v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-cross-lingual-information","repo_url":"https://github.com/rlitschk/UnsupCLIR","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"cross-lingual-information-retrieval","task_name":"Cross-Lingual Information Retrieval"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.00879","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}