{"url":"/dataset/arsen-20","name":"ArSen-20","full_name":null,"description_markdown":"Sentiment detection remains a pivotal task in natural language processing, yet its development in Arabic lags due to a scarcity of training materials compared to English. Addressing this gap, we present ArSen-20, a benchmark dataset tailored to propel Arabic sentiment detection forward. ArSen-20 comprises 20,000 professionally labeled tweets sourced from Twitter, focusing on the theme of COVID-19 and spanning the period from 2020 to 2023. Beyond tweet content, the dataset incorporates metadata associated with the user, enriching the contextual understanding. ArSen-20 offers a comprehensive resource to foster advancements in Arabic sentiment analysis and facilitate research in this critical domain.\r\n\r\nThe ArSen-20 dataset statistics:\r\n\r\n| Statistics   |   Num  | \r\n|:-------------:|:-----:|\r\n| Training set size | 16000 |\r\n| Validation set size| 2000 |\r\n| Testing set size | 2000 |\r\n| Neutral | 17262 |\r\n| Positive | 878 |\r\n| Negative | 1860 |\r\n\r\n## Features\r\nThe dataset has the following features:\r\n\r\n| Field   |  Type  |  Description  |\r\n|:-----------:| :--------: |:----------------: |\r\n| tweet id     | string     | The unique identifier of the requested Tweet.     |\r\n| label   | string     | Sentiment Classification of this tweet.     |\r\n| author id   | string    |The unique identifier of this user.     |\r\n| created_at  | data     | Creation time of the Tweet.    |\r\n| lang  | string     | Language of the Tweet, if detected by Twitter.    |\r\n| like_count  | int     |The number of likes on this tweet.|\r\n|quote_count  | int    | The number of times this tweet has been quoted.    |\r\n| reply_count   | int     | The number of replies to this tweet.    |\r\n| retweet_count| int    | The number of retweets to this tweet.    |\r\n| tweet   | string     | The actual UTF-8 text of the Tweet.    |\r\n|user_verified  | boolean     | Indicates if this user is a verified Twitter User.     |\r\n|followers_count  | int     |The number of followers of the author.     |\r\n| following_count  | int     | The number of following of the author.    |\r\n| tweet_count  | int     | Total number of tweets by the author.    |\r\n| listed_count | int     |The number of public lists that this user is a member of.    |\r\n|name | string     | The name of the user.    |\r\n| username   | string     | The Twitter screen name, handle, or alias.    |\r\n| user_created_at| data     | The UTC datetime that the user account was created.     |\r\n| description  | string     | The text of this user’s profile description (bio).     |\r\n\r\n## DownLoad\r\nYou can download the dataset from [here](https://github.com/123fangyang/ArSen-20/tree/main/data).\r\n\r\n+ ArSen-20_publish.csv - Contains all features.\r\n\r\n+ ArSen-20_id_only.csv - Contains only tweets and their author's id.\r\n\r\n\r\n## Citation\r\nIf you use this dataset in your research, please cite the following papers:\r\n```bibtex\r\n@inproceedings{fang2024arsen,\r\ntitle={ArSen-20: A New Benchmark for Arabic Sentiment Detection},\r\nauthor={Yang Fang and Cheng Xu},\r\nbooktitle={5th Workshop on African Natural Language Processing},\r\nyear={2024},\r\nurl={https://openreview.net/forum?id=GgsRUF5kJt}\r\n}\r\n```\r\n\r\n```bibtex\r\n@inproceedings{fang2024advancing,\r\n    title = \"Advancing {A}rabic Sentiment Analysis: {A}r{S}en Benchmark and the Improved Fuzzy Deep Hybrid Network\",\r\n    author = \"Fang, Yang  and\r\n      Xu, Cheng  and\r\n      Guan, Shuhao  and\r\n      Yan, Nan  and\r\n      Mei, Yuke\",\r\n    editor = \"Barak, Libby  and\r\n      Alikhani, Malihe\",\r\n    booktitle = \"Proceedings of the 28th Conference on Computational Natural Language Learning\",\r\n    month = nov,\r\n    year = \"2024\",\r\n    address = \"Miami, FL, USA\",\r\n    publisher = \"Association for Computational Linguistics\",\r\n    url = \"https://aclanthology.org/2024.conll-1.39\",\r\n    pages = \"507--516\",\r\n}\r\n```\r\n\r\n## contact\r\nIf you have any questions or comments about the dataset, please contact Yang Fang (20211209024@chnu.edu.cn).\r\n\r\nPotential cooperation in related fields is also welcome. :)","description_withheld":null,"homepage":"https://github.com/123fangyang/ArSen-20","introduced_date":"2024-04-11","introduced_date_note":null,"introduced_by":{"paper":"/paper/arsen-20-a-new-benchmark-for-arabic-sentiment","title":"ArSen-20: A New Benchmark for Arabic Sentiment Detection","first_author":"Yang Fang","url":null},"license":{"name":"Apache-2.0 license","url":"https://raw.githubusercontent.com/123fangyang/ArSen-20/main/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Sentiment Analysis","url":"/task/sentiment-analysis","datasets_with_task":"/datasets/task/sentiment-analysis"},{"name":"Arabic Sentiment Analysis","url":"/task/arabic-sentiment-analysis","datasets_with_task":"/datasets/task/arabic-sentiment-analysis"},{"name":"Twitter Sentiment Analysis","url":"/task/twitter-sentiment-analysis","datasets_with_task":"/datasets/task/twitter-sentiment-analysis"},{"name":"Sentiment Classification","url":"/task/sentiment-classification","datasets_with_task":"/datasets/task/sentiment-classification"}],"languages":[{"name":"Arabic","url":"/datasets/language/arabic"}],"variants":["ArSen-20"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}