{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/complexity-weighted-loss-and-diverse","title":"Complexity-Weighted Loss and Diverse Reranking for Sentence Simplification","arxiv_id":"1904.02767","date":"2019-04-04","proceeding":"NAACL 2019 6","authors":["Reno Kriz","João Sedoc","Marianna Apidianaki","Carolina Zheng","Gaurav Kumar","Eleni Miltsakaki","Chris Callison-Burch"],"abstract":"Sentence simplification is the task of rewriting texts so they are easier to\nunderstand. Recent research has applied sequence-to-sequence (Seq2Seq) models\nto this task, focusing largely on training-time improvements via reinforcement\nlearning and memory augmentation. One of the main problems with applying\ngeneric Seq2Seq models for simplification is that these models tend to copy\ndirectly from the original sentence, resulting in outputs that are relatively\nlong and complex. We aim to alleviate this issue through the use of two main\ntechniques. First, we incorporate content word complexities, as predicted with\na leveled word complexity model, into our loss function during training.\nSecond, we generate a large set of diverse candidate simplifications at test\ntime, and rerank these to promote fluency, adequacy, and simplicity. Here, we\nmeasure simplicity through a novel sentence complexity model. These extensions\nallow our models to perform competitively with state-of-the-art systems while\ngenerating simpler sentences. We report standard automatic and human evaluation\nmetrics.","url_abs":"http://arxiv.org/abs/1904.02767v1","url_pdf":"http://arxiv.org/pdf/1904.02767v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"complexity-weighted-loss-and-diverse","repo_url":"https://github.com/rekriz11/sockeye-recipes","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"complexity-weighted-loss-and-diverse","repo_url":"https://github.com/rekriz11/DeDiv","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reranking","task_name":"Reranking"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"text-simplification","task_name":"Text Simplification"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"seq2seq","method_name":"Seq2Seq"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-simplification-on-newsela","task":"Text Simplification","dataset":"Newsela","model":"S2S-Cluster-FA","rank_in_archive_order":4,"of":13,"metrics":{"BLEU":"19.55","SARI":"30.73"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.02767","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}