{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/posterior-regularized-reinforce-for-instance","title":"Posterior-regularized REINFORCE for Instance Selection in Distant Supervision","arxiv_id":"1904.08051","date":"2019-04-17","proceeding":"NAACL 2019 6","authors":["Qi Zhang","Siliang Tang","Xiang Ren","Fei Wu","ShiLiang Pu","Yueting Zhuang"],"abstract":"This paper provides a new way to improve the efficiency of the REINFORCE\ntraining process. We apply it to the task of instance selection in distant\nsupervision. Modeling the instance selection in one bag as a sequential\ndecision process, a reinforcement learning agent is trained to determine\nwhether an instance is valuable or not and construct a new bag with less noisy\ninstances. However unbiased methods, such as REINFORCE, could usually take much\ntime to train. This paper adopts posterior regularization (PR) to integrate\nsome domain-specific rules in instance selection using REINFORCE. As the\nexperiment results show, this method remarkably improves the performance of the\nrelation classifier trained on cleaned distant supervision dataset as well as\nthe efficiency of the REINFORCE training.","url_abs":"http://arxiv.org/abs/1904.08051v1","url_pdf":"http://arxiv.org/pdf/1904.08051v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"posterior-regularized-reinforce-for-instance","repo_url":"https://github.com/hitcszq/PRRLRE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"reinforce","method_name":"REINFORCE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}