{"url":"/sota/topic-coverage-on-topic-modeling-topic-1","task":{"name":"Topic coverage","url":"/task/topic-coverage","note":null},"dataset":{"name":"Topic modeling topic coverage dataset - bio","url":null},"category":"Natural Language Processing","categories":["Natural Language Processing"],"category_note":null,"description":"A prevalent use case of topic models is that of topic discovery.\r\nHowever, most of the topic model evaluation methods rely on abstract metrics such as perplexity or topic coherence. The topic coverage approach is to measure the models' performance by matching model-generated topics to a fixed set of reference topics - topics discovered by humans and represented in a machine-readable format. This way, the models are evaluated in the context of their use, by essentially simulating topic modeling in a fixed setting defined by a text collection and a set of reference topics.\r\nReference topics represent a ground truth that can be used to evaluate both topic models and other measures of model performance. This coverage approach enables large-scale automatic evaluation of existing and future topic models.","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["AuCDC","SupCov"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"AuCDC":null,"SupCov":null}},"counts":{"rows":2,"rows_with_code":2,"rows_with_paper_page":2,"rows_dated":2,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"NMF-200","metrics":{"AuCDC":"0.67","SupCov":"0.44"},"uses_additional_data":false,"paper_date":"2020-12-11","paper":"/paper/a-topic-coverage-approach-to-evaluation-of","paper_url":"https://arxiv.org/abs/2012.06274v3","paper_title":"A Topic Coverage Approach to Evaluation of Topic Models","code":"https://github.com/dkorenci/topic_coverage","n_code_links":1,"syntology":null},{"rank_in_archive_order":2,"model":"PYP","metrics":{"AuCDC":"0.56","SupCov":"0.23"},"uses_additional_data":false,"paper_date":"2020-12-11","paper":"/paper/a-topic-coverage-approach-to-evaluation-of","paper_url":"https://arxiv.org/abs/2012.06274v3","paper_title":"A Topic Coverage Approach to Evaluation of Topic Models","code":"https://github.com/dkorenci/topic_coverage","n_code_links":1,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":0,"rows_with_any_sample_ran":0,"distinct_papers_with_graph_line":0,"distinct_papers_with_any_sample_ran":0,"samples_over_distinct_papers":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}