{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/structure-to-property-chemical-element","title":"Structure to Property: Chemical Element Embeddings and a Deep Learning Approach for Accurate Prediction of Chemical Properties","arxiv_id":"2309.09355","date":"2023-09-17","proceeding":null,"authors":["Shokirbek Shermukhamedov","Dilorom Mamurjonova","Michael Probst"],"abstract":"We introduce the elEmBERT model for chemical classification tasks. It is based on deep learning techniques, such as a multilayer encoder architecture. We demonstrate the opportunities offered by our approach on sets of organic, inorganic and crystalline compounds. In particular, we developed and tested the model using the Matbench and Moleculenet benchmarks, which include crystal properties and drug design-related benchmarks. We also conduct an analysis of vector representations of chemical compounds, shedding light on the underlying patterns in structural data. Our model exhibits exceptional predictive capabilities and proves universally applicable to molecular and material datasets. For instance, on the Tox21 dataset, we achieved an average precision of 96%, surpassing the previously best result by 10%.","url_abs":"https://arxiv.org/abs/2309.09355v3","url_pdf":"https://arxiv.org/pdf/2309.09355v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"structure-to-property-chemical-element","repo_url":"https://github.com/dmamur/elembert","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"drug-design","task_name":"Drug Design"},{"task_slug":"drug-discovery","task_name":"Drug Discovery"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/drug-discovery-on-bace","task":"Drug Discovery","dataset":"BACE","model":"elEmBERT-V1","rank_in_archive_order":4,"of":6,"metrics":{"AUC":"0.856"},"uses_additional_data":false},{"leaderboard":"/sota/drug-discovery-on-bbbp","task":"Drug Discovery","dataset":"BBBP","model":"elEmBERT-V1","rank_in_archive_order":2,"of":4,"metrics":{"AUC":"0.905"},"uses_additional_data":false},{"leaderboard":"/sota/drug-discovery-on-sider","task":"Drug Discovery","dataset":"SIDER","model":"elEmBERT-V1","rank_in_archive_order":1,"of":4,"metrics":{"AUC":"0.778"},"uses_additional_data":false},{"leaderboard":"/sota/drug-discovery-on-tox21","task":"Drug Discovery","dataset":"Tox21","model":"elEmBERT-V1","rank_in_archive_order":1,"of":11,"metrics":{"AUC":"0.961"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}