{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/genhap-a-novel-computational-method-based-on","title":"GenHap: A Novel Computational Method Based on Genetic Algorithms for Haplotype Assembly","arxiv_id":"1812.07689","date":"2018-12-18","proceeding":null,"authors":[],"abstract":"The computational problem of inferring the full haplotype of a cell starting\nfrom read sequencing data is known as haplotype assembly, and consists in\nassigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly\none of the two chromosomes. Indeed, the knowledge of complete haplotypes is\ngenerally more informative than analyzing single SNPs and plays a fundamental\nrole in many medical applications. To reconstruct the two haplotypes, we\naddressed the weighted Minimum Error Correction (wMEC) problem, which is a\nsuccessful approach for haplotype assembly. This NP-hard problem consists in\ncomputing the two haplotypes that partition the sequencing reads into two\ndisjoint sub-sets, with the least number of corrections to the SNP values. To\nthis aim, we propose here GenHap, a novel computational method for haplotype\nassembly based on Genetic Algorithms, yielding optimal solutions by means of a\nglobal search process. In order to evaluate the effectiveness of our approach,\nwe run GenHap on two synthetic (yet realistic) datasets, based on the Roche/454\nand PacBio RS II sequencing technologies. We compared the performance of GenHap\nagainst HapCol, an efficient state-of-the-art algorithm for haplotype phasing.\nOur results show that GenHap always obtains high accuracy solutions (in terms\nof haplotype error rate), and is up to 4x faster than HapCol in the case of\nRoche/454 instances and up to 20x faster when compared on the PacBio RS II\ndataset. Finally, we assessed the performance of GenHap on two different real\ndatasets. Future-generation sequencing technologies, producing longer reads\nwith higher coverage, can highly benefit from GenHap, thanks to its capability\nof efficiently solving large instances of the haplotype assembly problem.","url_abs":"http://arxiv.org/abs/1812.07689v1","url_pdf":"http://arxiv.org/pdf/1812.07689v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"genhap-a-novel-computational-method-based-on","repo_url":"https://github.com/andrea-tango/GenHap","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}