{"url":"/dataset/sd-eval","name":"SD-Eval","full_name":null,"description_markdown":"Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.\r\nThis comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction.\r\nChat-Oriented Large Language Models (LLMs), known for their general-purpose assistance capabilities, have evolved to handle multi-modal inputs, including speech.\r\nAlthough these models can be adept at recognizing and analyzing speech, they often fall short of generating appropriate responses.\r\nWe argue that this is due to the lack of principles on task definition and model development, which requires open-source datasets and metrics suitable for model evaluation.\r\nTo bridge the gap, we present SD-Eval, a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation.\r\nSD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data. The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound.\r\nTo assess the SD-Eval benchmark dataset, we implement three different models and construct a training set following a process similar to that of SD-Eval. The training set contains 1,052.72 hours of speech data and 724.4k utterances. \r\nWe also conduct a comprehensive evaluation using objective evaluation methods (e.g. BLEU and ROUGE), subjective evaluations and LLM-based metrics for the generated responses.\r\nModels conditioned with paralinguistic and environmental information outperform their counterparts in both objective and subjective measures.\r\nMoreover, experiments demonstrate that LLM-based metrics show a higher correlation with human evaluation compared to traditional metrics.\r\nWe open-source SD-Eval at https://github.com/amphionspace/SD-Eval.","description_withheld":null,"homepage":"https://github.com/amphionspace/SD-Eval","introduced_date":"2024-06-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/sd-eval-a-benchmark-dataset-for-spoken","title":"SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words","first_author":"Junyi Ao","url":null},"license":{"name":"CC BY-NC 4.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Audio","url":"/datasets/modality/audio"},{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Spoken Dialogue Systems","url":"/task/spoken-dialogue-systems","datasets_with_task":"/datasets/task/spoken-dialogue-systems"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SD-Eval"],"data_loaders":[],"num_papers_in_archive":9,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}