{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/independent-vector-analysis-for-data-fusion","title":"Independent Vector Analysis for Data Fusion Prior to Molecular Property Prediction with Machine Learning","arxiv_id":"1811.00628","date":"2018-11-01","proceeding":null,"authors":["Zois Boukouvalas","Daniel C. Elton","Peter W. Chung","Mark D. Fuge"],"abstract":"Due to its high computational speed and accuracy compared to ab-initio\nquantum chemistry and forcefield modeling, the prediction of molecular\nproperties using machine learning has received great attention in the fields of\nmaterials design and drug discovery. A main ingredient required for machine\nlearning is a training dataset consisting of molecular features\\textemdash for\nexample fingerprint bits, chemical descriptors, etc. that adequately\ncharacterize the corresponding molecules. However, choosing features for any\napplication is highly non-trivial. No \"universal\" method for feature selection\nexists. In this work, we propose a data fusion framework that uses Independent\nVector Analysis to exploit underlying complementary information contained in\ndifferent molecular featurization methods, bringing us a step closer to\nautomated feature generation. Our approach takes an arbitrary number of\nindividual feature vectors and automatically generates a single, compact (low\ndimensional) set of molecular features that can be used to enhance the\nprediction performance of regression models. At the same time our methodology\nretains the possibility of interpreting the generated features to discover\nrelationships between molecular structures and properties. We demonstrate this\non the QM7b dataset for the prediction of several properties such as\natomization energy, polarizability, frontier orbital eigenvalues, ionization\npotential, electron affinity, and excitation energies. In addition, we show how\nour method helps improve the prediction of experimental binding affinities for\na set of human BACE-1 inhibitors. To encourage more widespread use of IVA we\nhave developed the PyIVA Python package, an open source code which is available\nfor download on Github.","url_abs":"http://arxiv.org/abs/1811.00628v1","url_pdf":"http://arxiv.org/pdf/1811.00628v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"independent-vector-analysis-for-data-fusion","repo_url":"https://github.com/zoisboukouvalas/pyiva","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"drug-discovery","task_name":"Drug Discovery"},{"task_slug":"molecular-property-prediction","task_name":"Molecular Property Prediction"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"property-prediction","task_name":"Property Prediction"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}