{"url":"/dataset/pcd-polish-cyberbullying-dataset","name":"PCD (Polish Cyberbullying Dataset)","full_name":null,"description_markdown":"The **Polish Cyberbullying Dataset** is a valuable resource for studying harmful online phenomena, specifically **cyberbullying** and **hate speech** in the Polish language. Let's delve into the details:\r\n\r\n1. **Dataset Overview**:\r\n   - The dataset contains **tweets** that have been **annotated** with labels indicating whether they fall into the category of **cyberbullying/harmful** or **non-cyberbullying/non-harmful** content.\r\n   - These tweets were collected from **publicly available Twitter discussions**.\r\n   - Minimal preprocessing has been applied to the tweets, primarily focusing on cases where information about a private individual is revealed to the public.\r\n\r\n2. **Purpose and Significance**:\r\n   - The lack of research on cyberbullying detection in the Polish language prompted the creation of this dataset.\r\n   - Researchers aim to **automatically detect harmful content** on the internet and report it to service providers for further analysis and removal.\r\n   - The dataset facilitates the development and evaluation of **classification methods** for identifying cyberbullying-related narratives.\r\n\r\n3. **Tasks Available**:\r\n   - Users can experiment with their **classification algorithms** to determine whether an internet entry (tweet) is part of a cyberbullying narrative or not.\r\n   - The primary goal is to classify tweets into either **cyberbullying/harmful** or **non-cyberbullying/non-harmful** categories, achieving high precision, recall, balanced F-score, and accuracy.\r\n   - An additional sub-task focuses on differentiating between various types of harmful information, such as **cyberbullying** or **hate speech**.\r\n\r\n4. **Example Format of Tweets**:\r\n   - Each tweet is represented as follows:\r\n     ```\r\n     \"contents\", class (1=harmful, 0=non-harmful, etc.)\r\n     ```\r\n\r\n5. **Accessing the Dataset**:\r\n   - The dataset is available on **GitHub** under the repository named **\"cyberbullying-Polish\"**¹.\r\n   - It includes annotated labels, scoring scripts, and raw tweets for research purposes.\r\n\r\nThis dataset contributes significantly to addressing the growing problem of harmful online behavior, providing researchers with a foundation for developing automatic cyberbullying detection methods in the Polish language ².\r\n\r\nSource: Conversation with Bing, 3/16/2024\r\n(1) GitHub - ptaszynski/cyberbullying-Polish: This dataset contains tweets .... https://github.com/ptaszynski/cyberbullying-Polish.\r\n(2) Expert-Annotated Dataset to Study Cyberbullying in Polish Language - MDPI. https://www.mdpi.com/2306-5729/9/1/1.\r\n(3) Data | Free Full-Text | Expert-Annotated Dataset to Study Cyberbullying .... https://www.mdpi.com/2306-5729/9/1/1/notes.\r\n(4) undefined. http://poleval.pl/tasks/task6.","description_withheld":null,"homepage":"https://github.com/ptaszynski/cyberbullying-Polish","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["PCD (Polish Cyberbullying Dataset)"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}