{"url":"/dataset/topic-modeling-topic-coverage-dataset","name":"Topic modeling topic coverage dataset","full_name":null,"description_markdown":"A prevalent use case of topic models is that of topic discovery.\r\nHowever, most of the topic model evaluation methods rely on abstract metrics such as perplexity or topic coherence.\r\nThe topic coverage approach is to measure the models' performance by matching model-generated topics to topics discovered by humans.\r\nThis way, the models are evaluated in the context of their use, by essentially simulating\r\ntopic modeling in a fixed setting defined by a text collection and a set of reference topics.\r\n\r\nReference topics represent a ground truth that can be used to evaluate both topic models and other measures of model performance.\r\nThe coverage approach enables large-scale automatic evaluation of both existing and future topic models.\r\n\r\nThe topic coverage dataset consists of two text collections and two sets of reference topics.\r\nThese two sub-datasets correspond to two domains (news text and biological text) \r\nwhere topic models are used for topic discovery in large text collections.\r\nThe reference topics consist of model-generated topics inspected, selected, and curated by humans.\r\n\r\nEach dataset contains a corpus of preprocessed (tokenized) texts and a set of reference topics, \r\neach represented by a list of words and text documents.\r\nThe dataset details, including the instruction for the use of the data and supporting code, are here:\r\nhttps://github.com/dkorenci/topic_coverage/blob/main/data.readme.txt\r\n\r\nThe coverage measures that can be used to evaluate topic models are described in the accompanying paper, \r\nwhereas the code and the instructions can be found in the github repo.","description_withheld":null,"homepage":"https://github.com/dkorenci/topic_coverage","introduced_date":"2021-08-31","introduced_date_note":null,"introduced_by":{"paper":"/paper/a-topic-coverage-approach-to-evaluation-of","title":"A Topic Coverage Approach to Evaluation of Topic Models","first_author":"Damir Korenčić","url":null},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Topic coverage","url":"/task/topic-coverage","datasets_with_task":"/datasets/task/topic-coverage"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Topic modeling topic coverage dataset"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/topic-coverage-on-topic-modeling-topic-2","task":"Topic coverage","dataset_variant":"Topic modeling topic coverage dataset","rows":1,"metrics":["Spearman Correlation"],"first_row_in_archive_order":{"model":"AuCDC","paper":"/paper/a-topic-coverage-approach-to-evaluation-of","metrics":{"Spearman Correlation":"0.95"},"code_links":[{"title":"dkorenci/topic_coverage","url":"https://github.com/dkorenci/topic_coverage"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/a-topic-coverage-approach-to-evaluation-of","title":"A Topic Coverage Approach to Evaluation of Topic Models","date":"2020-12-11","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}