{"url":"/dataset/social-media-messages-for-early-cyberattack","name":"Social Media Messages for Early Cyberattack Detection on Blockchain","full_name":null,"description_markdown":"# ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity\r\n🔐 **12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain**\r\n\r\n## Dataset Statistics\r\n\r\n| **Category** | **Samples** | **Description** |\r\n|--------------|-------------|-----------------|\r\n| **Cyberattack** | 6,941 | Early warning signals and indicators of cyberattacks |\r\n| **General** | 4,507 | Regular blockchain discussions (non-security related) |\r\n\r\n## Dataset Structure\r\n\r\nEach entry in the dataset contains:\r\n- `message_id`: Unique identifier for each message\r\n- `content`: The text content of the social media message\r\n- `topic`: Classification label (\"cyberattack\" or \"general\")\r\n\r\n## Performance \r\n[Gemma-2b-it](https://huggingface.co/google/gemma-2b-it) fine-tuned on this dataset:\r\n- Achieves a Brier score of 0.16 using only synthetic data in our social media threat detection task\r\n- Shows competitive performance on this specific task when compared to general-purpose models known for capabilities in cybersecurity tasks, like [granite-3.2-2b-instruct](https://huggingface.co/ibm-granite/granite-3.2-2b-instruct), and cybersecurity-focused LLMs trained on [Primus](https://huggingface.co/papers/2502.11191)\r\n- Demonstrates promising results for smaller models on this specific task, with our best hybrid model achieving an F1 score of 0.81 on our blockchain threat detection test set, though GPT-4o maintains superior overall accuracy (0.84) and calibration (Brier 0.10)\r\n\r\n## Attack Type Distribution\r\n\r\n| **Attack Vectors** | **Seed Examples** | \r\n|-------------------|------------------------|\r\n| Social Engineering & Phishing | Credential theft, wallet phishing | \r\n| Smart Contract Exploits | Token claim vulnerabilities, flash loans | \r\n| Exchange Security Breaches | Hot wallet compromises, key theft |\r\n| DeFi Protocol Attacks | Liquidity pool manipulation, bridge exploits | \r\n\r\nFor more about Cyberattack Vectors, read [Attack Vectors Wiki](https://dn.institute/research/cyberattacks/wiki/)\r\n\r\n## Citation\r\n\r\n```bibtex\r\n@misc{razmyslovich2025eltexframeworkdomaindrivensynthetic,\r\n      title={ELTEX: A Framework for Domain-Driven Synthetic Data Generation}, \r\n      author={Arina Razmyslovich and Kseniia Murasheva and Sofia Sedlova and Julien Capitaine and Eugene Dmitriev},\r\n      year={2025},\r\n      eprint={2503.15055},\r\n      archivePrefix={arXiv},\r\n      primaryClass={cs.CL},\r\n      url={https://arxiv.org/abs/2503.15055}, \r\n}\r\n```","description_withheld":null,"homepage":"https://huggingface.co/datasets/dn-institute/cyberattack-blockchain-synth","introduced_date":"2025-03-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/eltex-a-framework-for-domain-driven-synthetic","title":"ELTEX: A Framework for Domain-Driven Synthetic Data Generation","first_author":"Arina Razmyslovich","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Social Media Messages for Early Cyberattack Detection on Blockchain"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}