{"url":"/dataset/wmt-2021-ge-ez-amharic","name":"WMT 2021 Ge'ez-Amharic","full_name":null,"description_markdown":"**WMT 2021 Ge'ez-Amharic** is a Ge'ez-Amharic  dataset prepared for NMT tasks of the 6th Workshop on NLP at Debre Berhan University, Ethiopia. The corpus has been collected from:\r\n\r\n* Ethiopian Orthodox Church old bible (from ethiopianorthodox.org), Anaphora, praise of St. Virgin Mary, praise of Lord Jesus and other Church's books.\r\n* Ge'ez teaching books,\r\n* Websites and other internet sources such as www.geez.org, www.debelo.org, \r\n\r\nThe Dataset has about 15454 parallel Ge'ez and Amharic sentences for training, 1001 parallel sentences for testing and 1001 parallel sentences  for validation.","description_withheld":null,"homepage":"https://github.com/Amdework21/Geez-Amharic-DS","introduced_date":"2017-06-12","introduced_date_note":null,"introduced_by":null,"license":{"name":"Unknown","url":"https://github.com/Amdework21/Geez-Amharic-DS"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Machine Translation","url":"/task/machine-translation","datasets_with_task":"/datasets/task/machine-translation"}],"languages":[{"name":"Amharic","url":"/datasets/language/amharic"},{"name":"Geez","url":"/datasets/language/geez"}],"variants":["WMT 2021 Ge'ez-Amharic"],"data_loaders":[{"repo":"https://github.com/Amdework21/Geez-Amharic-DS","url":"https://github.com/Amdework21/Geez-Amharic-DS","frameworks":[]}],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}