Papers › CoRAL: a Context-aware Croatian Abusive Language Dataset

CoRAL: a Context-aware Croatian Abusive Language Dataset

11 Nov 2022arXiv:2211.06053archive 2025-07-28

Ravi Shekhar, Mladen Karan, Matthew Purver

In light of unprecedented increases in the popularity of the internet and social media, comment moderation has never been a more relevant task. Semi-automated comment moderation systems greatly aid human moderators by either automatically classifying the examples or allowing the moderators to prioritize which comments to consider first. However, the concept of inappropriate content is often subjective, and such content can be conveyed in many subtle and indirect ways. In this work, we propose CoRAL -- a language and culturally aware Croatian Abusive dataset covering phenomena of implicitness and reliance on local and global context. We show experimentally that current models degrade when comments are not explicit and further degrade when language skill and context knowledge are required to interpret the comment.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Abusive Language

Datasets

Introduced by this paper, per the archive.

CoRAL dataset

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

AWARECORAL

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections