{"url":"/dataset/web-forum-52","name":"WEB-FORUM-52","full_name":"WEB-FORUM-52 gold standard","description_markdown":"The WEB-FORUM-52 gold standard comprises (i) 13 web forums from the health domain, (ii) 15 forums obtained from a Wikipedia list of popular forums (https://en.wikipedia.org/wiki/List_of_Internet_forums), (iii) 13 forums mentioned on a list of popular German Web forums (https://www.beliebte-foren.de), (iv) nine forums obtained from WPressBlog (https://www.wpressblog.com/free-forum-posting-sites-list/) and (v) two additional forums. For most forums two web pages (from different threads) were used and stored together with gold standard annotations that have been manually created by domain experts and describe the post text, post date, post user and direct URL to the post.","description_withheld":null,"homepage":"https://github.com/fhgr/harvest","introduced_date":"2020-10-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/harvest-an-open-source-toolkit-for-extracting","title":"Harvest -- An Open Source Toolkit for Extracting Posts and Post Metadata from Web Forums","first_author":"Albert Weichselbraun","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"German","url":"/datasets/language/german"}],"variants":["WEB-FORUM-52"],"data_loaders":[{"repo":"https://github.com/fhgr/harvest","url":"https://github.com/fhgr/harvest","frameworks":[]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}