{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skip-gram-language-modeling-using-sparse-non","title":"Skip-gram Language Modeling Using Sparse Non-negative Matrix Probability Estimation","arxiv_id":"1412.1454","date":"2014-12-03","proceeding":null,"authors":["Noam Shazeer","Joris Pelemans","Ciprian Chelba"],"abstract":"We present a novel family of language model (LM) estimation techniques named\nSparse Non-negative Matrix (SNM) estimation. A first set of experiments\nempirically evaluating it on the One Billion Word Benchmark shows that SNM\n$n$-gram LMs perform almost as well as the well-established Kneser-Ney (KN)\nmodels. When using skip-gram features the models are able to match the\nstate-of-the-art recurrent neural network (RNN) LMs; combining the two modeling\ntechniques yields the best known result on the benchmark. The computational\nadvantages of SNM over both maximum entropy and RNN LM estimation are probably\nits main strength, promising an approach that has the same flexibility in\ncombining arbitrary features effectively and yet should scale to very large\namounts of data as gracefully as $n$-gram LMs do.","url_abs":"http://arxiv.org/abs/1412.1454v2","url_pdf":"http://arxiv.org/pdf/1412.1454v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/language-modelling-on-one-billion-word","task":"Language Modelling","dataset":"One Billion Word","model":"Sparse Non-Negative","rank_in_archive_order":25,"of":27,"metrics":{"Number of params":"33B","PPL":"52.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1412.1454","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}