{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bias-and-high-dimensional-adjustment-in","title":"Bias and high-dimensional adjustment in observational studies of peer effects","arxiv_id":"1706.04692","date":"2017-06-14","proceeding":null,"authors":["Dean Eckles","Eytan Bakshy"],"abstract":"Peer effects, in which the behavior of an individual is affected by the\nbehavior of their peers, are posited by multiple theories in the social\nsciences. Other processes can also produce behaviors that are correlated in\nnetworks and groups, thereby generating debate about the credibility of\nobservational (i.e. nonexperimental) studies of peer effects. Randomized field\nexperiments that identify peer effects, however, are often expensive or\ninfeasible. Thus, many studies of peer effects use observational data, and\nprior evaluations of causal inference methods for adjusting observational data\nto estimate peer effects have lacked an experimental \"gold standard\" for\ncomparison. Here we show, in the context of information and media diffusion on\nFacebook, that high-dimensional adjustment of a nonexperimental control group\n(677 million observations) using propensity score models produces estimates of\npeer effects statistically indistinguishable from those from using a large\nrandomized experiment (220 million observations). Naive observational\nestimators overstate peer effects by 320% and commonly used variables (e.g.,\ndemographics) offer little bias reduction, but adjusting for a measure of prior\nbehaviors closely related to the focal behavior reduces bias by 91%.\nHigh-dimensional models adjusting for over 3,700 past behaviors provide\nadditional bias reduction, such that the full model reduces bias by over 97%.\nThis experimental evaluation demonstrates that detailed records of individuals'\npast behavior can improve studies of social influence, information diffusion,\nand imitation; these results are encouraging for the credibility of some\nstudies but also cautionary for studies of rare or new behaviors. More\ngenerally, these results show how large, high-dimensional data sets and\nstatistical learning techniques can be used to improve causal inference in the\nbehavioral sciences.","url_abs":"http://arxiv.org/abs/1706.04692v1","url_pdf":"http://arxiv.org/pdf/1706.04692v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bias-and-high-dimensional-adjustment-in","repo_url":"https://github.com/fghjorth/vkme18","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"causal-inference","task_name":"Causal Inference"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[{"method_slug":"causal-inference","method_name":"Causal inference"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}