{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/smiles-enumeration-as-data-augmentation-for","title":"SMILES Enumeration as Data Augmentation for Neural Network Modeling of Molecules","arxiv_id":"1703.07076","date":"2017-03-21","proceeding":null,"authors":["Esben Jannik Bjerrum"],"abstract":"Simplified Molecular Input Line Entry System (SMILES) is a single line text\nrepresentation of a unique molecule. One molecule can however have multiple\nSMILES strings, which is a reason that canonical SMILES have been defined,\nwhich ensures a one to one correspondence between SMILES string and molecule.\nHere the fact that multiple SMILES represent the same molecule is explored as a\ntechnique for data augmentation of a molecular QSAR dataset modeled by a long\nshort term memory (LSTM) cell based neural network. The augmented dataset was\n130 times bigger than the original. The network trained with the augmented\ndataset shows better performance on a test set when compared to a model built\nwith only one canonical SMILES string per molecule. The correlation coefficient\nR2 on the test set was improved from 0.56 to 0.66 when using SMILES\nenumeration, and the root mean square error (RMS) likewise fell from 0.62 to\n0.55. The technique also works in the prediction phase. By taking the average\nper molecule of the predictions for the enumerated SMILES a further improvement\nto a correlation coefficient of 0.68 and a RMS of 0.52 was found.","url_abs":"http://arxiv.org/abs/1703.07076v2","url_pdf":"http://arxiv.org/pdf/1703.07076v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"smiles-enumeration-as-data-augmentation-for","repo_url":"https://github.com/Ebjerrum/SMILES-enumeration","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"smiles-enumeration-as-data-augmentation-for","repo_url":"https://github.com/EBjerrum/molvecgen","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"smiles-enumeration-as-data-augmentation-for","repo_url":"https://github.com/MolecularAI/pysmilesutils","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"smiles-enumeration-as-data-augmentation-for","repo_url":"https://github.com/lantunes/chemgrams","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.07076","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}