{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-better-understanding-of-artifacts-in","title":"Towards Better Understanding of Artifacts in Variant Calling from High-Coverage Samples","arxiv_id":"1404.0929","date":"2014-04-03","proceeding":null,"authors":["Heng Li"],"abstract":"Motivation: Whole-genome high-coverage sequencing has been widely used for\npersonal and cancer genomics as well as in various research areas. However, in\nthe lack of an unbiased whole-genome truth set, the global error rate of\nvariant calls and the leading causal artifacts still remain unclear even given\nthe great efforts in the evaluation of variant calling methods.\n  Results: We made ten SNP and INDEL call sets with two read mappers and five\nvariant callers, both on a haploid human genome and a diploid genome at a\nsimilar coverage. By investigating false heterozygous calls in the haploid\ngenome, we identified the erroneous realignment in low-complexity regions and\nthe incomplete reference genome with respect to the sample as the two major\nsources of errors, which press for continued improvements in these two areas.\nWe estimated that the error rate of raw genotype calls is as high as 1 in\n10-15kb, but the error rate of post-filtered calls is reduced to 1 in 100-200kb\nwithout significant compromise on the sensitivity.\n  Availability: BWA-MEM alignment: http://bit.ly/1g8XqRt; Scripts:\nhttps://github.com/lh3/varcmp; Additional data:\nhttps://figshare.com/articles/Towards_better_understanding_of_artifacts_in_variating_calling_from_high_coverage_samples/981073","url_abs":"http://arxiv.org/abs/1404.0929v2","url_pdf":"http://arxiv.org/pdf/1404.0929v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-better-understanding-of-artifacts-in","repo_url":"https://github.com/lh3/varcmp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"towards-better-understanding-of-artifacts-in","repo_url":"https://github.com/elixir-no-nels/rbFlow-Germline","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"articles","task_name":"Articles"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}