{"url":"/dataset/languagenet","name":"LanguageNet","full_name":null,"description_markdown":"The LanguageNet (English) is a collection of sentence level paraphrases from Twitter by linking tweets through shared URLs. This corpus is the largest up to date with 51,524 human annotated sentence pairs: 42200 for training and 9324 for testing. It can grow 30,000 new sentential paraphrases per month with ~70% precision. Now we have 1-year data available: 2,869,657 candidate pairs!","description_withheld":null,"homepage":"https://languagenet.github.io/","introduced_date":"2018-08-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/languagenet-learning-to-find-sense-relevant","title":"LanguageNet: Learning to Find Sense Relevant Example Sentences","first_author":"Shang-Chien Cheng","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["LanguageNet"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}