{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/uncovering-divergent-linguistic-information","title":"Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation","arxiv_id":"1809.02094","date":"2018-09-06","proceeding":"CONLL 2018 10","authors":["Mikel Artetxe","Gorka Labaka","Iñigo Lopez-Gazpio","Eneko Agirre"],"abstract":"Following the recent success of word embeddings, it has been argued that\nthere is no such thing as an ideal representation for words, as different\nmodels tend to capture divergent and often mutually incompatible aspects like\nsemantics/syntax and similarity/relatedness. In this paper, we show that each\nembedding model captures more information than directly apparent. A linear\ntransformation that adjusts the similarity order of the model without any\nexternal resource can tailor it to achieve better results in those aspects,\nproviding a new perspective on how embeddings encode divergent linguistic\ninformation. In addition, we explore the relation between intrinsic and\nextrinsic evaluation, as the effect of our transformations in downstream tasks\nis higher for unsupervised systems than for supervised ones.","url_abs":"http://arxiv.org/abs/1809.02094v1","url_pdf":"http://arxiv.org/pdf/1809.02094v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"uncovering-divergent-linguistic-information","repo_url":"https://github.com/artetxem/uncovec","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"uncovering-divergent-linguistic-information","repo_url":"https://github.com/lgazpio/DAM_STS","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1809.02094","atlas_url":"https://app.syntology.ai/?focus=1809.02094","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}