Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Vijay Murugesan

Predicting Aqueous Solubility of Organic Molecules Using Deep Learning Models with Varied Molecular Representations

May 27, 2021

Gihan Panapitiya, Michael Girard, Aaron Hollas, Vijay Murugesan, Wei Wang, Emily Saldanha

Figure 1 for Predicting Aqueous Solubility of Organic Molecules Using Deep Learning Models with Varied Molecular Representations

Figure 2 for Predicting Aqueous Solubility of Organic Molecules Using Deep Learning Models with Varied Molecular Representations

Figure 3 for Predicting Aqueous Solubility of Organic Molecules Using Deep Learning Models with Varied Molecular Representations

Figure 4 for Predicting Aqueous Solubility of Organic Molecules Using Deep Learning Models with Varied Molecular Representations

Abstract:Determining the aqueous solubility of molecules is a vital step in many pharmaceutical, environmental, and energy storage applications. Despite efforts made over decades, there are still challenges associated with developing a solubility prediction model with satisfactory accuracy for many of these applications. The goal of this study is to develop a general model capable of predicting the solubility of a broad range of organic molecules. Using the largest currently available solubility dataset, we implement deep learning-based models to predict solubility from molecular structure and explore several different molecular representations including molecular descriptors, simplified molecular-input line-entry system (SMILES) strings, molecular graphs, and three-dimensional (3D) atomic coordinates using four different neural network architectures - fully connected neural networks (FCNNs), recurrent neural networks (RNNs), graph neural networks (GNNs), and SchNet. We find that models using molecular descriptors achieve the best performance, with GNN models also achieving good performance. We perform extensive error analysis to understand the molecular properties that influence model performance, perform feature analysis to understand which information about molecular structure is most valuable for prediction, and perform a transfer learning and data size study to understand the impact of data availability on model performance.

Via

Access Paper or Ask Questions