{"url":"/dataset/linkso","name":"LinkSO","full_name":null,"description_markdown":"The LinkSO dataset is a resource for learning to retrieve similar question-answer pairs on Stack Overflow. It consists of three datasets corresponding to three popular programming languages (Python, Java, JavaScript), 690K question pairs, and 26K linked question pairs (i.e., positive examples). The dataset was extracted from Stack Overflow's data dump in April 2018 and was cleaned and pre-processed to remove non-ASCII characters, email addresses, URLs, and code blocks. The dataset was designed to help propose new models, such as neural network models, to improve community-based question-answer retrieval in the software engineering domain.","description_withheld":null,"homepage":"https://sites.google.com/view/linkso","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["LinkSO"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}