{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dreaddit-a-reddit-dataset-for-stress-analysis-1","title":"Dreaddit: A Reddit Dataset for Stress Analysis in Social Media","arxiv_id":"1911.00133","date":"2019-10-31","proceeding":"WS 2019 11","authors":["Elsbeth Turcan","Kathleen McKeown"],"abstract":"Stress is a nigh-universal human experience, particularly in the online world. While stress can be a motivator, too much stress is associated with many negative health outcomes, making its identification useful across a range of domains. However, existing computational research typically only studies stress in domains such as speech, or in short genres such as Twitter. We present Dreaddit, a new text corpus of lengthy multi-domain social media data for the identification of stress. Our dataset consists of 190K posts from five different categories of Reddit communities; we additionally label 3.5K total segments taken from 3K posts using Amazon Mechanical Turk. We present preliminary supervised learning methods for identifying stress, both neural and traditional, and analyze the complexity and diversity of the data and characteristics of each category.","url_abs":"https://arxiv.org/abs/1911.00133v1","url_pdf":"https://arxiv.org/pdf/1911.00133v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dreaddit-a-reddit-dataset-for-stress-analysis-1","repo_url":"https://github.com/gillian850413/Insight_Stress_Analysis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"}],"methods":[],"datasets_introduced":[{"slug":"dreaddit","name":"Dreaddit","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1911.00133","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1911.00133"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/gillian850413/Insight_Stress_Analysis","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":3},"by_repo_kind":{"listed":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e318ebf8e16f7d3c","entry":"get_corpus","repo":"gillian850413/Insight_Stress_Analysis","repo_kind":"listed","path":"script/tfidf.py","file_url":"https://github.com/gillian850413/Insight_Stress_Analysis/blob/HEAD/script/tfidf.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e318ebf8e16f7d3c"}},{"code_sha256_prefix":"066426ed269ffa7d","entry":"get_tfidf_vector","repo":"gillian850413/Insight_Stress_Analysis","repo_kind":"listed","path":"script/tfidf.py","file_url":"https://github.com/gillian850413/Insight_Stress_Analysis/blob/HEAD/script/tfidf.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"066426ed269ffa7d"}},{"code_sha256_prefix":"bb36331f29df2e85","entry":"get_top_tfidf_features","repo":"gillian850413/Insight_Stress_Analysis","repo_kind":"listed","path":"script/tfidf.py","file_url":"https://github.com/gillian850413/Insight_Stress_Analysis/blob/HEAD/script/tfidf.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bb36331f29df2e85"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}