{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/swivel-improving-embeddings-by-noticing-whats","title":"Swivel: Improving Embeddings by Noticing What's Missing","arxiv_id":"1602.02215","date":"2016-02-06","proceeding":null,"authors":["Noam Shazeer","Ryan Doherty","Colin Evans","Chris Waterson"],"abstract":"We present Submatrix-wise Vector Embedding Learner (Swivel), a method for\ngenerating low-dimensional feature embeddings from a feature co-occurrence\nmatrix. Swivel performs approximate factorization of the point-wise mutual\ninformation matrix via stochastic gradient descent. It uses a piecewise loss\nwith special handling for unobserved co-occurrences, and thus makes use of all\nthe information in the matrix. While this requires computation proportional to\nthe size of the entire matrix, we make use of vectorized multiplication to\nprocess thousands of rows and columns at once to compute millions of predicted\nvalues. Furthermore, we partition the matrix into shards in order to\nparallelize the computation across many nodes. This approach results in more\naccurate embeddings than can be achieved with methods that consider only\nobserved co-occurrences, and can scale to much larger corpora than can be\nhandled with sampling methods.","url_abs":"http://arxiv.org/abs/1602.02215v1","url_pdf":"http://arxiv.org/pdf/1602.02215v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"swivel-improving-embeddings-by-noticing-whats","repo_url":"https://github.com/IBM/MAX-Word-Embedding-Generator","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"swivel-improving-embeddings-by-noticing-whats","repo_url":"https://github.com/tensorflow/models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"swivel-improving-embeddings-by-noticing-whats","repo_url":"https://github.com/tensorflow/models/tree/master/research/swivel","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1602.02215","atlas_url":"https://app.syntology.ai/?focus=1602.02215","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}