{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-multi-level-threats-in-telegram","title":"Exploring Multi-Level Threats in Telegram Data with AI-Human Annotation: A Preliminary Study","arxiv_id":null,"date":"2023-12-15","proceeding":"2023 22nd IEEE International Conference on Machine Learning and Applications (ICMLA) 2023 12","authors":["Kamalakkannan Ravi","Adan Ernesto Vela","Elizabeth Jenaway","Steven Windisch"],"abstract":"This research addresses the crucial challenge of effectively measuring threats in social media comments targeting voting, public officials, and institutions in the United States. Our understanding of these online threats and their links to real-world risks is limited, making it difficult to assess their seriousness. To overcome these limitations, we propose a comprehensive threat level scale from 0 to 5 and collect a dataset of 1.3 million Telegram responses for developing and rigorously testing these threat levels. Additionally, we explore OpenAI-human annotation to efficiently label this vast dataset. Our innovative two-step transfer learning approach initially employs a pre-existing, pre-trained model for labeling, followed by expert validation. Next, we use the AI-annotated samples to develop independent models, and expert annotators verify their predictions. Notably, our findings demonstrate that the GPT-2 model, despite its fewer annotated training set, performs comparably to OpenAI's anno-tations, showcasing its potential for cost-effective threat detection with more annotated samples. With the long-term objective of establishing continuous threat-level monitoring, we identify the strengths and limitations of our current approach and propose a roadmap for enhancing threat detection.","url_abs":"https://ieeexplore.ieee.org/abstract/document/10459792","url_pdf":"https://ieeexplore.ieee.org/abstract/document/10459792","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"violence-and-weaponized-violence-detection","task_name":"Violence and Weaponized Violence Detection"}],"methods":[{"method_slug":"awd-lstm","method_name":"AWD-LSTM"},{"method_slug":"activation-regularization","method_name":"Activation Regularization"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"attention-dropout","method_name":"Attention Dropout"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"cosine-annealing","method_name":"Cosine Annealing"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"discriminative-fine-tuning","method_name":"Discriminative Fine-Tuning"},{"method_slug":"dropconnect","method_name":"DropConnect"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"embedding-dropout","method_name":"Embedding Dropout"},{"method_slug":"gpt-2","method_name":"GPT-2"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"linear-warmup-with-cosine-annealing","method_name":"Linear Warmup With Cosine Annealing"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"svm","method_name":"SVM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"slanted-triangular-learning-rates","method_name":"Slanted Triangular Learning Rates"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"},{"method_slug":"temporal-activation-regularization","method_name":"Temporal Activation Regularization"},{"method_slug":"ulmfit","method_name":"ULMFiT"},{"method_slug":"variational-dropout","method_name":"Variational Dropout"},{"method_slug":"weight-decay","method_name":"Weight Decay"},{"method_slug":"weight-tying","method_name":"Weight Tying"},{"method_slug":"fasttext","method_name":"fastText"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-classification-on-threatgram-101-extreme","task":"Text Classification","dataset":"ThreatGram 101 - Extreme Telegram Data","model":"GPT-2","rank_in_archive_order":1,"of":4,"metrics":{"weighted-F1 score":"66.2"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-threatgram-101-extreme","task":"Text Classification","dataset":"ThreatGram 101 - Extreme Telegram Data","model":"SVM","rank_in_archive_order":2,"of":4,"metrics":{"weighted-F1 score":"64.3"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-threatgram-101-extreme","task":"Text Classification","dataset":"ThreatGram 101 - Extreme Telegram Data","model":"fastText","rank_in_archive_order":3,"of":4,"metrics":{"weighted-F1 score":"60.2"},"uses_additional_data":false},{"leaderboard":"/sota/text-classification-on-threatgram-101-extreme","task":"Text Classification","dataset":"ThreatGram 101 - Extreme Telegram Data","model":"ULMFit","rank_in_archive_order":4,"of":4,"metrics":{"weighted-F1 score":"55.7"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}