{"url":"/dataset/qald-9-plus","name":"QALD-9-Plus","full_name":null,"description_markdown":"# QALD-9-Plus Dataset Description\r\n\r\n[QALD-9-Plus](https://github.com/Perevalov/qald_9_plus) is the dataset for Knowledge Graph Question Answering (KGQA) based on well-known [QALD-9](https://github.com/ag-sc/QALD/tree/master/9/data).\r\n\r\nQALD-9-Plus enables to train and test KGQA systems over DBpedia and Wikidata using questions in 9 different languages: English, German, Russian, French, Armenian, Belarusian, Lithuanian, Bashkir, and Ukrainian.\r\n\r\nSome of the questions have several alternative writings in particular languages which enables to evaluate the robustness of KGQA systems and train paraphrasing models.\r\n\r\nAs the questions' translations were provided by native speakers, they are considered as \"gold standard\", therefore, machine translation tools can be trained and evaluated on the dataset.\r\n\r\n# Dataset Statistics\r\n\r\n|       |  en |  de | fr |  ru  |  uk |  lt |  be |  ba | hy | # questions DBpedia | # questions Wikidata |\r\n|-------|:---:|:---:|:--:|:----:|:---:|:---:|:---:|:---:|:--:|:-----------:|:-----------:|\r\n| Train | 408 | 543 | 260 | 1203 | 447 | 468 | 441 | 284 | 80 |     408     |     371     |\r\n| Test  | 150 | 176 | 26 |  348 | 176 | 186 | 155 | 117 | 20 |     150     |     136     |\r\n\r\nGiven the numbers, it is obvious that some of the languages are covered more than once i.e., there is more than one translation for a particular question.\r\nFor example, there are 1203 Russian translations available while only 408 unique questions exist in the training subset (i.e., 2.9 Russian translations per one question).\r\nThe availability of such parallel corpora enables the researchers, developers and other dataset users to address the paraphrasing task.\r\n\r\n[cc-by]: http://creativecommons.org/licenses/by/4.0/\r\n[cc-by-image]: https://i.creativecommons.org/l/by/4.0/88x31.png\r\n[cc-by-shield]: https://img.shields.io/badge/License-CC%20BY%204.0-lightgrey.svg","description_withheld":null,"homepage":"https://github.com/Perevalov/qald_9_plus","introduced_date":"2022-01-31","introduced_date_note":null,"introduced_by":{"paper":"/paper/qald-9-plus-a-multilingual-dataset-for","title":"QALD-9-plus: A Multilingual Dataset for Question Answering over DBpedia and Wikidata Translated by Native Speakers","first_author":"Aleksandr Perevalov","url":null},"license":{"name":"Creative Commons Attribution 4.0 International License","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Question Answering","url":"/task/question-answering","datasets_with_task":"/datasets/task/question-answering"},{"name":"Knowledge Base Question Answering","url":"/task/knowledge-base-question-answering","datasets_with_task":"/datasets/task/knowledge-base-question-answering"},{"name":"Cross-Lingual Question Answering","url":"/task/cross-lingual-question-answering","datasets_with_task":"/datasets/task/cross-lingual-question-answering"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"French","url":"/datasets/language/french"},{"name":"German","url":"/datasets/language/german"},{"name":"Multilingual","url":"/datasets/language/multilingual"},{"name":"Russian","url":"/datasets/language/russian"},{"name":"Armenian","url":"/datasets/language/armenian"},{"name":"Belarusian","url":"/datasets/language/belarusian"},{"name":"Lithuanian","url":"/datasets/language/lithuanian"},{"name":"Ukrainian","url":"/datasets/language/ukrainian"},{"name":"Bashkir","url":"/datasets/language/bashkir"}],"variants":["QALD-9-Plus"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/knowledge-base-question-answering-on-qald-9","task":"Knowledge Base Question Answering","dataset_variant":"QALD-9-Plus","rows":12,"metrics":["Macro F1"],"first_row_in_archive_order":{"model":"QAnswer-Wikidata-English","paper":"/paper/qald-9-plus-a-multilingual-dataset-for","metrics":{"Macro F1":"0.4459"},"code_links":[{"title":"perevalov/qald_9_plus","url":"https://github.com/perevalov/qald_9_plus"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/qald-9-plus-a-multilingual-dataset-for","title":"QALD-9-plus: A Multilingual Dataset for Question Answering over DBpedia and Wikidata Translated by Native Speakers","date":"2022-01-31","rows_on_this_dataset":12,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}