{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-continuously-growing-dataset-of-sentential","title":"A Continuously Growing Dataset of Sentential Paraphrases","arxiv_id":"1708.00391","date":"2017-08-01","proceeding":"EMNLP 2017 9","authors":["Wuwei Lan","Siyu Qiu","Hua He","Wei Xu"],"abstract":"A major challenge in paraphrase research is the lack of parallel corpora. In\nthis paper, we present a new method to collect large-scale sentential\nparaphrases from Twitter by linking tweets through shared URLs. The main\nadvantage of our method is its simplicity, as it gets rid of the classifier or\nhuman in the loop needed to select data before annotation and subsequent\napplication of paraphrase identification algorithms in the previous work. We\npresent the largest human-labeled paraphrase corpus to date of 51,524 sentence\npairs and the first cross-domain benchmarking for automatic paraphrase\nidentification. In addition, we show that more than 30,000 new sentential\nparaphrases can be easily and continuously captured every month at ~70%\nprecision, and demonstrate their utility for downstream NLP tasks through\nphrasal paraphrase extraction. We make our code and data freely available.","url_abs":"http://arxiv.org/abs/1708.00391v1","url_pdf":"http://arxiv.org/pdf/1708.00391v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"paraphrase-identification","task_name":"Paraphrase Identification"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[{"slug":"twitter-news-url-corpus","name":"TURL","full_name":"Twitter News URL Corpus"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.00391","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}