Papers › Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the...

Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People's Republic of China

14 Jun 2021arXiv:2106.07544archive 2025-07-28

Rong-Ching Chang, Chun-Ming Lai, Kai-Lai Chang, Chu-Hsing Lin

The digital media, identified as computational propaganda provides a pathway for propaganda to expand its reach without limit. State-backed propaganda aims to shape the audiences' cognition toward entities in favor of a certain political party or authority. Furthermore, it has become part of modern information warfare used in order to gain an advantage over opponents. Most of the current studies focus on using machine learning, quantitative, and qualitative methods to distinguish if a certain piece of information on social media is propaganda. Mainly conducted on English content, but very little research addresses Chinese Mandarin content. From propaganda detection, we want to go one step further to provide more fine-grained information on propaganda techniques that are applied. In this research, we aim to bridge the information gap by providing a multi-labeled propaganda techniques dataset in Mandarin based on a state-backed information operation dataset provided by Twitter. In addition to presenting the dataset, we apply a multi-label text classification using fine-tuned BERT. Potentially this could help future research in detecting state-backed propaganda online especially in a cross-lingual context and cross platforms identity consolidation.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

annabechang/Propaganda_Tech_Twitter_PRC officialmentioned on GitHubMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Multi Label Text ClassificationMulti-Label Text ClassificationPropaganda detectionText Classificationtext-classification

Datasets

Introduced by this paper, per the archive.

Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People's Republic of China

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Multi-Label Text Classification Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People's Republic of China Bert 1:1 Accuracy 0.80352 #1 of 1 Archive leaderboard report
Multi-Label Text Classification Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People's Republic of China Bert F1 - macro 0.20803 #1 of 1 Archive leaderboard report
Multi-Label Text Classification Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People's Republic of China Bert Micro F1 0.85431 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections