{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/protvec-a-continuous-distributed","title":"ProtVec: A Continuous Distributed Representation of Biological Sequences","arxiv_id":"1503.05140","date":"2015-03-17","proceeding":null,"authors":["Ehsaneddin Asgari","Mohammad R. K. Mofrad"],"abstract":"We introduce a new representation and feature extraction method for\nbiological sequences. Named bio-vectors (BioVec) to refer to biological\nsequences in general with protein-vectors (ProtVec) for proteins (amino-acid\nsequences) and gene-vectors (GeneVec) for gene sequences, this representation\ncan be widely used in applications of deep learning in proteomics and genomics.\nIn the present paper, we focus on protein-vectors that can be utilized in a\nwide array of bioinformatics investigations such as family classification,\nprotein visualization, structure prediction, disordered protein identification,\nand protein-protein interaction prediction. In this method, we adopt artificial\nneural network approaches and represent a protein sequence with a single dense\nn-dimensional vector. To evaluate this method, we apply it in classification of\n324,018 protein sequences obtained from Swiss-Prot belonging to 7,027 protein\nfamilies, where an average family classification accuracy of 93%+-0.06% is\nobtained, outperforming existing family classification methods. In addition, we\nuse ProtVec representation to predict disordered proteins from structured\nproteins. Two databases of disordered sequences are used: the DisProt database\nas well as a database featuring the disordered regions of nucleoporins rich\nwith phenylalanine-glycine repeats (FG-Nups). Using support vector machine\nclassifiers, FG-Nup sequences are distinguished from structured protein\nsequences found in Protein Data Bank (PDB) with a 99.8% accuracy, and\nunstructured DisProt sequences are differentiated from structured DisProt\nsequences with 100.0% accuracy. These results indicate that by only providing\nsequence data for various proteins into this model, accurate information about\nprotein structure can be determined.","url_abs":"http://arxiv.org/abs/1503.05140v2","url_pdf":"http://arxiv.org/pdf/1503.05140v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"protvec-a-continuous-distributed","repo_url":"https://github.com/ehsanasgari/Deep-Proteomics","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}