{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/decoding-decoders-finding-optimal","title":"Decoding Decoders: Finding Optimal Representation Spaces for Unsupervised Similarity Tasks","arxiv_id":"1805.03435","date":"2018-05-09","proceeding":"ICLR 2018 1","authors":["Vitalii Zhelezniak","Dan Busbridge","April Shen","Samuel L. Smith","Nils Y. Hammerla"],"abstract":"Experimental evidence indicates that simple models outperform complex deep\nnetworks on many unsupervised similarity tasks. We provide a simple yet\nrigorous explanation for this behaviour by introducing the concept of an\noptimal representation space, in which semantically close symbols are mapped to\nrepresentations that are close under a similarity measure induced by the\nmodel's objective function. In addition, we present a straightforward procedure\nthat, without any retraining or architectural modifications, allows deep\nrecurrent models to perform equally well (and sometimes better) when compared\nto shallow models. To validate our analysis, we conduct a set of consistent\nempirical evaluations and introduce several new sentence embedding models in\nthe process. Even though this work is presented within the context of natural\nlanguage processing, the insights are readily applicable to other domains that\nrely on distributed representations for transfer tasks.","url_abs":"http://arxiv.org/abs/1805.03435v1","url_pdf":"http://arxiv.org/pdf/1805.03435v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"decoding-decoders-finding-optimal","repo_url":"https://github.com/Babylonpartners/decoding-decoders","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentence-embedding","task_name":"Sentence Embedding"},{"task_slug":"sentence-embedding-1","task_name":"Sentence-Embedding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}