{"url":"/dataset/arabic-tod","name":"Arabic-ToD","full_name":"Arabic-ToD: Arabic Task Oriented Dialogue dataset","description_markdown":"The Arabic-TOD dataset is based on the BiToD dataset.\r\nOf the 3,689 BiToD-English dialogues, 1,500 dialogues (30,000 utterances) were translated into Arabic.\r\nWe translated the task-related keywords such as cuisine, dietary restrictions, and price-level for the\r\nrestaurant domain, price-level for the hotel domain, type, and price-level for the attraction domain, day,\r\nweather, and city for the weather domain. We keep the rest of values without translation, like hotels’ and\r\nrestaurants’ names, locations, and addresses. These values are real  entities in Hong Kong city (literals),\r\nand most of them contain Chinese words written in English, therefore they have not been translated. According to\r\nthe slot-values in the Arabic-TOD dataset, we used the slots names as they are in English and translated their\r\ncorresponding values, except the entities in Hong Kong city since the Arabic-TOD dataset supports codeswitching. \r\n\r\nWe did not translate the 'UserTask' for all dialogues, since it is not important in developing the system.\r\nIt is just as a summarization of the dialogue contents.","description_withheld":null,"homepage":"https://github.com/En-J-A/Arabic-TOD","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[{"name":"Dialogue Generation","url":"/task/dialogue-generation","datasets_with_task":"/datasets/task/dialogue-generation"},{"name":"Task-Oriented Dialogue Systems","url":"/task/task-oriented-dialogue-systems","datasets_with_task":"/datasets/task/task-oriented-dialogue-systems"},{"name":"End-To-End Dialogue Modelling","url":"/task/end-to-end-dialogue-modelling","datasets_with_task":"/datasets/task/end-to-end-dialogue-modelling"},{"name":"Conversational Response Generation","url":"/task/conversational-response-generation","datasets_with_task":"/datasets/task/conversational-response-generation"}],"languages":[],"variants":["Arabic-ToD"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}