{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-training-of-neural-trans-dimensional","title":"Improved training of neural trans-dimensional random field language models with dynamic noise-contrastive estimation","arxiv_id":"1807.00993","date":"2018-07-03","proceeding":null,"authors":["Bin Wang","Zhijian Ou"],"abstract":"A new whole-sentence language model - neural trans-dimensional random field\nlanguage model (neural TRF LM), where sentences are modeled as a collection of\nrandom fields, and the potential function is defined by a neural network, has\nbeen introduced and successfully trained by noise-contrastive estimation (NCE).\nIn this paper, we extend NCE and propose dynamic noise-contrastive estimation\n(DNCE) to solve the two problems observed in NCE training. First, a dynamic\nnoise distribution is introduced and trained simultaneously to converge to the\ndata distribution. This helps to significantly cut down the noise sample number\nused in NCE and reduce the training cost. Second, DNCE discriminates between\nsentences generated from the noise distribution and sentences generated from\nthe interpolation of the data distribution and the noise distribution. This\nalleviates the overfitting problem caused by the sparseness of the training\nset. With DNCE, we can successfully and efficiently train neural TRF LMs on\nlarge corpus (about 0.8 billion words) with large vocabulary (about 568 K\nwords). Neural TRF LMs perform as good as LSTM LMs with less parameters and\nbeing 5x~114x faster in rescoring sentences. Interpolating neural TRF LMs with\nLSTM LMs and n-gram LMs can further reduce the error rates.","url_abs":"http://arxiv.org/abs/1807.00993v1","url_pdf":"http://arxiv.org/pdf/1807.00993v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improved-training-of-neural-trans-dimensional","repo_url":"https://github.com/wbengine/TRF-NN-Tensorflow","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1807.00993","atlas_url":"https://app.syntology.ai/?focus=1807.00993","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}