{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-pairwise-coevolutionary-models-capture","title":"How pairwise coevolutionary models capture the collective residue variability in proteins","arxiv_id":"1801.04184","date":"2018-01-12","proceeding":null,"authors":[],"abstract":"Global coevolutionary models of homologous protein families, as constructed\nby direct coupling analysis (DCA), have recently gained popularity in\nparticular due to their capacity to accurately predict residue-residue contacts\nfrom sequence information alone, and thereby to facilitate tertiary and\nquaternary protein structure prediction. More recently, they have also been\nused to predict fitness effects of amino-acid substitutions in proteins, and to\npredict evolutionary conserved protein-protein interactions. These models are\nbased on two currently unjustified hypotheses: (a) correlations in the\namino-acid usage of different positions are resulting collectively from\nnetworks of direct couplings; and (b) pairwise couplings are sufficient to\ncapture the amino-acid variability. Here we propose a highly precise inference\nscheme based on Boltzmann-machine learning, which allows us to systematically\naddress these hypotheses. We show how correlations are built up in a highly\ncollective way by a large number of coupling paths, which are based on the\nprotein's three-dimensional structure. We further find that pairwise\ncoevolutionary models capture the collective residue variability across\nhomologous proteins even for quantities which are not imposed by the inference\nprocedure, like three-residue correlations, the clustered structure of protein\nfamilies in sequence space or the sequence distances between homologs. These\nfindings strongly suggest that pairwise coevolutionary models are actually\nsufficient to accurately capture the residue variability in homologous protein\nfamilies.","url_abs":"http://arxiv.org/abs/1801.04184v1","url_pdf":"http://arxiv.org/pdf/1801.04184v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-pairwise-coevolutionary-models-capture","repo_url":"https://github.com/matteofigliuzzi/bmDCA","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"protein-structure-prediction","task_name":"Protein Structure Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1801.04184","atlas_url":"https://app.syntology.ai/?focus=1801.04184","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}