{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/from-paraphrase-database-to-compositional","title":"From Paraphrase Database to Compositional Paraphrase Model and Back","arxiv_id":"1506.03487","date":"2015-06-10","proceeding":"TACL 2015 1","authors":["John Wieting","Mohit Bansal","Kevin Gimpel","Karen Livescu","Dan Roth"],"abstract":"The Paraphrase Database (PPDB; Ganitkevitch et al., 2013) is an extensive\nsemantic resource, consisting of a list of phrase pairs with (heuristic)\nconfidence estimates. However, it is still unclear how it can best be used, due\nto the heuristic nature of the confidences and its necessarily incomplete\ncoverage. We propose models to leverage the phrase pairs from the PPDB to build\nparametric paraphrase models that score paraphrase pairs more accurately than\nthe PPDB's internal scores while simultaneously improving its coverage. They\nallow for learning phrase embeddings as well as improved word embeddings.\nMoreover, we introduce two new, manually annotated datasets to evaluate\nshort-phrase paraphrasing models. Using our paraphrase model trained using\nPPDB, we achieve state-of-the-art results on standard word and bigram\nsimilarity tasks and beat strong baselines on our new short phrase paraphrase\ntasks.","url_abs":"http://arxiv.org/abs/1506.03487v2","url_pdf":"http://arxiv.org/pdf/1506.03487v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"from-paraphrase-database-to-compositional","repo_url":"https://github.com/madcpt/pretrained-embeddings-toolkit","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}