{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/protein-identification-with-deep-learning","title":"Protein identification with deep learning: from abc to xyz","arxiv_id":"1710.02765","date":"2017-10-08","proceeding":null,"authors":["Ngoc Hieu Tran","Zachariah Levine","Lei Xin","Baozhen Shan","Ming Li"],"abstract":"Proteins are the main workhorses of biological functions in a cell, a tissue,\nor an organism. Identification and quantification of proteins in a given\nsample, e.g. a cell type under normal/disease conditions, are fundamental tasks\nfor the understanding of human health and disease. In this paper, we present\nDeepNovo, a deep learning-based tool to address the problem of protein\nidentification from tandem mass spectrometry data. The idea was first proposed\nin the context of de novo peptide sequencing [1] in which convolutional neural\nnetworks and recurrent neural networks were applied to predict the amino acid\nsequence of a peptide from its spectrum, a similar task to generating a caption\nfrom an image. We further develop DeepNovo to perform sequence database search,\nthe main technique for peptide identification that greatly benefits from\nnumerous existing protein databases. We combine two modules de novo sequencing\nand database search into a single deep learning framework for peptide\nidentification, and integrate de Bruijn graph assembly technique to offer a\ncomplete solution to reconstruct protein sequences from tandem mass\nspectrometry data. This paper describes a comprehensive protocol of DeepNovo\nfor protein identification, including training neural network models, dynamic\nprogramming search, database querying, estimation of false discovery rate, and\nde Bruijn graph assembly. Training and testing data, model implementations, and\ncomprehensive tutorials in form of IPython notebooks are available in our\nGitHub repository (https://github.com/nh2tran/DeepNovo).","url_abs":"http://arxiv.org/abs/1710.02765v1","url_pdf":"http://arxiv.org/pdf/1710.02765v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"protein-identification-with-deep-learning","repo_url":"https://github.com/nh2tran/DeepNovo","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"de-novo-peptide-sequencing","task_name":"de novo peptide sequencing"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}