Papers › RICo: Reddit ideological communities
RICo: Reddit ideological communities
Kamalakkannan Ravi, Adan Ernesto Vela
The main objective of our research is to gain a comprehensive understanding of the relationship between language usage within different communities and delineating the ideological narratives. We focus specifically on utilizing Natural Language Processing techniques to identify underlying narratives in the coded or suggestive language employed by non-normative communities associated with targeted violence. Earlier studies addressed the detection of ideological affiliation through surveys, user studies, and a limited number based on the content of text articles, which still require label curation. Previous work addressed label curation by using ideological subreddits (r/Liberal and r/Conservative for Liberal and Conservative classes) to label the articles shared on those subreddits according to their prescribed ideologies, albeit with a limited dataset. Building upon previous work, we use subreddit ideologies to categorize shared articles. In addition to the conservative and liberal classes, we introduce a new category called “Restricted” which encompasses text articles shared in subreddits that are restricted, privatized, or banned, such as r/TheDonald. The “Restricted” class encompasses posts tied to violence, regardless of conservative or liberal affiliations. Additionally, we augment our dataset with text articles from self-identified subreddits like r/progressive and r/askaconservative for the liberal and conservative classes, respectively. This results in an expanded dataset of 377,144 text articles, consisting of 72,488 liberal, 79,573 conservative, and 225,083 restricted class articles. Our goal is to analyze language variances in different ideological communities, investigate keyword relevance in labeling article orientations, especially in unseen cases (922,522 text articles), and delve into radicalized communities, conducting thorough analysis and interpretation of the results.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| News Classification | Reddit Ideological and Extreme Bias Dataset | SVM | weighted-F1 score | 79.1 | #1 of 7 | Archive leaderboard | report |
| News Classification | Reddit Ideological and Extreme Bias Dataset | ULMFit | weighted-F1 score | 76.8 | #2 of 7 | Archive leaderboard | report |
| News Classification | Reddit Ideological and Extreme Bias Dataset | Longformer | weighted-F1 score | 76.47 | #3 of 7 | Archive leaderboard | report |
| News Classification | Reddit Ideological and Extreme Bias Dataset | GPT-2 | weighted-F1 score | 76.43 | #4 of 7 | Archive leaderboard | report |
| News Classification | Reddit Ideological and Extreme Bias Dataset | RoBERTa | weighted-F1 score | 75.2 | #5 of 7 | Archive leaderboard | report |
| News Classification | Reddit Ideological and Extreme Bias Dataset | fastText | weighted-F1 score | 73.44 | #6 of 7 | Archive leaderboard | report |
| News Classification | Reddit Ideological and Extreme Bias Dataset | LightGBM | weighted-F1 score | 73.06 | #7 of 7 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections