Home › Datasets › language › Arabic

Arabic datasets

archive 2025-07-28

109 datasets carry the language tag "Arabic", ordered by the archive's paper count. Page 2 of 3: 48 shown of 109. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Arabic datasets 49–96 of 109

CATT (CATT Arabic Diacritization Benchmark Dataset)
The CATT benchmark dataset comprises 742 sentences, which were scraped from an internet news source in 2023.
3 papers · 1 benchmark
DAWT (Densely Annotated Wikipedia Texts)
The DAWT dataset consists of Densely Annotated Wikipedia Texts across multiple languages.
3 papers · 0 benchmarks
DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
This paper introduces the RGB Arabic Alphabet Sign Language (AASL) dataset.
3 papers · 0 benchmarks
Stanceosaurus is a corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.
3 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
ANETAC (Arabic Named Entity Transliteration and Classification)
An English-Arabic named entity transliteration and classification dataset built from freely available parallel translation corpora.
2 papers · 0 benchmarks
Natural Language Inference processes pairs of sentences to extract their semantic relations.
2 papers · 0 benchmarks
Sentiment detection remains a pivotal task in natural language processing, yet its development in Arabic lags due to a scarcity of training materials compared to English.
2 papers · 0 benchmarks
AraCOVID19-MFH (AraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News and Hate Speech Detection Dataset)
AraCOVID19-MFH is a manually annotated multi-label Arabic COVID-19 fake news and hate speech detection dataset.
2 papers · 0 benchmarks
Kaleidoscope (Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation)
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
QDAT data set contains 1500 WAV files along with sound files stored on Excel CSV file format.
2 papers · 0 benchmarks
A human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
2 papers · 0 benchmarks
Semantic Question Similarity in Arabic (NSURL-2019 Shared Task 8: Semantic Question Similarity in Arabic)
NSURL-2019 Shared Task 8: Semantic Question Similarity in Arabic This dataset contains 11,997 pairs of questions in MSA Arabic that are assigned either a label of 0, for no semantic similarity, or 1 otherwise.
2 papers · 0 benchmarks
A new text effects dataset with 141,081 text effect/glyph pairs in total.
2 papers · 0 benchmarks
AMFDS (Arabic Multi-Fonts Dataset)
Arabic Multi Fonts Dataset A multi-word multi-font Arabic word-image dataset.
1 paper · 0 benchmarks
AQL-22 (Archive Query Log)
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
AjwaOrMedjool (AjwaOrMedjool: a binary balanced dataset to teach machine learning‏)
The dataset contains three subsets: 1- a dataset containing hand-crafted features to classify two types of organic dates (Ajwa or Medjool); 2- a dataset containing tabular data with features created automatically using deep learning to…
1 paper · 0 benchmarks
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
1 paper · 0 benchmarks
ArEEGChars, the first EEG dataset for Arabic characters, consists of high-quality recordings for 31 unique characters from 30 participants (21 males and 9 females) using the Epoc X 14-channel device.
1 paper · 0 benchmarks
ArEEGWords dataset is a novel EEG dataset recorded from 22 participants with mean age of 22 years (5 female, 17 male) using a 14-channel Emotiv Epoc X device.
1 paper · 0 benchmarks
Sentiment analysis is pivotal in Natural Language Processing for understanding opinions and emotions in text.
1 paper · 0 benchmarks
ArVoice (ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis)
We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration,…
1 paper · 0 benchmarks
AraCovid19-SSD is a manually annotated Arabic COVID-19 sarcasm and sentiment detection dataset containing 5,162 tweets.
1 paper · 0 benchmarks
This dataset is designed to help train simple machine learning models that serve educational and research purposes in the speech recognition domain, mainly for keyword spotting tasks.
1 paper · 0 benchmarks
The dataset covers three types of medical interactions in both English and Arabic: - Multiple-choice question answering (MCQA), focusing on specialized medical knowledge.
1 paper · 0 benchmarks
Box-Jenkins (Box-Jenkins Gas Furnace Problem)
Box-Jenkins gas furnace, a well-known time series forecasting problem
1 paper · 0 benchmarks
Calliar is a dataset for Arabic calligraphy.
1 paper · 0 benchmarks
DIGITal (Digitally Generated Numerals)
Digitally Generated Numerals (DIGITal) Description The Digitally Generated Numerals (DIGITal) dataset consists of 100,000 image pairs representing digits from 0 to 9.
1 paper · 0 benchmarks
Contains 350 tweets with more than 8,000 words including 3,000 unique words written in Egyptian dialect.
1 paper · 0 benchmarks
The ExaASC dataset is a dataset for Target-based Stance Detection in the Arabic Language that contains different types of targets like persons, entities and events.
1 paper · 0 benchmarks
GLARE is an Arabic Apps Reviews dataset collected from Saudi Google PlayStore.
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
HARD (Hotel Arabic-Reviews Dataset)
The Hotel Arabic-Reviews Dataset (HARD) contains 93700 hotel reviews in Arabic language.
1 paper · 1 benchmark
KHATT (KFUPM Handwritten Arabic TexT Database)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
LeT-Mi (Levantine Twitter dataset for Misogynistic language)
Levantine Twitter dataset for Misogynistic language (LeT-Mi) is an Arabic Levantine Twitter dataset for misogynistic language to be the first benchmark dataset for Arabic misogyny.
1 paper · 0 benchmarks
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
The AASL-Clear dataset is a collection of RGB images featuring Arabic alphabet sign Language gestures with backgrounds removed.
1 paper · 1 benchmark
This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh).
1 paper · 0 benchmarks
Poly-FEVER is a multilingual fact verification benchmark designed to evaluate hallucination detection in large language models (LLMs).
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
PsOCR (Pashto OCR Dataset)
PsOCR is a large-scale synthetic dataset for Optical Character Recognition in low-resource Pashto language.
1 paper · 0 benchmarks
RGB Arabic Alphabet Sign Language (AASL) dataset
1 paper · 1 benchmark
TQBA++ (Tiny QA Benchmark++)
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.