{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-document-embeddings-by-predicting-n","title":"Learning Document Embeddings by Predicting N-grams for Sentiment Classification of Long Movie Reviews","arxiv_id":"1512.08183","date":"2015-12-27","proceeding":null,"authors":["Bofang Li","Tao Liu","Xiaoyong Du","Deyuan Zhang","Zhe Zhao"],"abstract":"Despite the loss of semantic information, bag-of-ngram based methods still\nachieve state-of-the-art results for tasks such as sentiment classification of\nlong movie reviews. Many document embeddings methods have been proposed to\ncapture semantics, but they still can't outperform bag-of-ngram based methods\non this task. In this paper, we modify the architecture of the recently\nproposed Paragraph Vector, allowing it to learn document vectors by predicting\nnot only words, but n-gram features as well. Our model is able to capture both\nsemantics and word order in documents while keeping the expressive power of\nlearned vectors. Experimental results on IMDB movie review dataset shows that\nour model outperforms previous deep learning models and bag-of-ngram based\nmodels due to the above advantages. More robust results are also obtained when\nour model is combined with other models. The source code of our model will be\nalso published together with this paper.","url_abs":"http://arxiv.org/abs/1512.08183v5","url_pdf":"http://arxiv.org/pdf/1512.08183v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-document-embeddings-by-predicting-n","repo_url":"https://github.com/libofang/DV-ngram","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"},{"task_slug":"sentiment-classification","task_name":"Sentiment Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}