{"url":"/dataset/openmeva","name":"OpenMEVA","full_name":null,"description_markdown":"OpenMEVA is a benchmark for evaluating open-ended story generation metrics. OpenMEVA provides a comprehensive test suite to assess the capabilities of metrics, including (a) the correlation with human judgments, (b) the generalization to different model outputs and datasets, (c) the ability to judge story coherence, and (d) the robustness to perturbations. To this end, OpenMEVA includes both manually annotated stories and auto-constructed test examples.\r\n\r\nSource: [OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics](https://arxiv.org/pdf/2105.08920v1.pdf)\r\n\r\nImage source: [OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics](https://arxiv.org/pdf/2105.08920v1.pdf)","description_withheld":null,"homepage":"https://github.com/thu-coai/OpenMEVA","introduced_date":"2021-05-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/openmeva-a-benchmark-for-evaluating-open","title":"OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics","first_author":"Jian Guan","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[],"variants":["OpenMEVA"],"data_loaders":[],"num_papers_in_archive":10,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}