{"url":"/dataset/ccpm","name":"CCPM","full_name":"Chinese Classical Poetry Matching","description_markdown":"**Introduction**\r\n\r\nCCPM is a large Chinese classical poetry matching dataset that can be used for poetry matching, understanding and translation.\r\n\r\nThe main task of this dataset is: given a description in modern Chinese, the model is supposed to select one line of Chinese classical poetry from four candidates that semantically match the given description most.\r\n\r\n**Size**\r\n\r\nIt contains 27,218 instances in total, which are split into training (21,778), validation (2,720) and test (2,720) sets.\r\n\r\n**Format**\r\n\r\nEach instance is composed of translation (the description in modern Chinese, a string), choice (four candidate lines of Chinese classical poetry, a list) and answer (the index of the correct line, an integer between 0 and 3).\r\n\r\nSource: [https://github.com/THUNLP-AIPoet/CCPM](https://github.com/THUNLP-AIPoet/CCPM)","description_withheld":null,"homepage":"https://github.com/THUNLP-AIPoet/CCPM","introduced_date":"2021-06-03","introduced_date_note":null,"introduced_by":{"paper":"/paper/ccpm-a-chinese-classical-poetry-matching","title":"CCPM: A Chinese Classical Poetry Matching Dataset","first_author":"Wenhao Li","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["CCPM"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}