{"url":"/dataset/bpersona-chat","name":"BPersona-chat","full_name":null,"description_markdown":"BPersona-chat is an evaluation dataset based on the English multiturn chat corpus Persona-chat and the Japanese multiturn chat corpus JPersona-chat.\r\n\r\nEach chat was performed between two crowd workers assuming artificial personas. The speakers discuss a given personality trait, including but not limited to self-introduction, hobby, and others. (Notice that they are not translations of each other.)\r\n\r\nChats are translated into Japanese/English by professional translators, a low-quality machine translation model A and a high-quality machine translation model B.\r\n\r\nTranslations are evaluated by crowdworkers as either good or bad, depending on the correctness and coherence.\r\n\r\nEach chat is included in one .xlsx file with the following structure:\r\n\r\nperson - the speaker on the current utterance,\r\nsource - the utterance in the source language,\r\ntranslation - the translation in the target language,\r\nevaluation: is this a good translation? - the evaluation of the translation's quality,\r\ny - the current translation is a correct translation of the source utterance,\r\nn - the current translation is an erroneous translation of the source utterance.","description_withheld":null,"homepage":"https://github.com/cl-tohoku/BPersona-chat","introduced_date":"2022-12-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/chat-translation-error-detection-for","title":"Chat Translation Error Detection for Assisting Cross-lingual Communications","first_author":"Yunmeng Li","url":null},"license":{"name":"CC BY-NC 4.0","url":"https://creativecommons.org/licenses/by-nc/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Machine Translation","url":"/task/machine-translation","datasets_with_task":"/datasets/task/machine-translation"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"Japanese","url":"/datasets/language/japanese"}],"variants":["BPersona-chat"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}