International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1703191

1703191 Vol 5 · Issue 8 Download Paper

Data Analytics ? Computer Modelling of Metabolic Rates

Sunday Tarekakpo Odobai Nazifi Lawal Bashir

Subject area: Science,Engineering and Technology  ·  Area of research: Analytics

Abstract

Artificial Neural Networks (ANNs) and Multiple Linear Regression (MLR) based Quantitative Structure-Activity Relationships (QSARs) models were developed to predict enzymatic activities, that is, the Michaelis-Menten constant (Km) and the maximum reaction rate (Vmax) for reactions involving the biotransformation of xenobiotics, catalysed by three classes of enzymes present in the mammalian livers. The enzymes we have studied here are alcohol dehydrogenase (ADH), aldehyde dehydrogenase (ALDH), and Flavin-containing monooxygenase (FMO). Data for enzymatic constants were collected from the literature and the computation of potential predictors was done for all xenobiotics to include for hundreds of molecular descriptors. The best predictor variables were selected (maximum of seven and a minimum of two descriptors) using the Microsoft excel correlation function for each enzyme class. Each dataset was divided into three sets, the divisions were training, cross-validation, and test sets in the ratio of 70%, 15%, and 15% respectively for both the ANNs and the MLR models to build the QSARs. The MATLAB programming language was employed to implement the writing and running of the learning algorithms. The predictive strengths of the models were assessed through the correlation of their predictions relative to the target outcomes for the three divisions and the mean square errors were computed, after fitting the resulting models with the entire dataset for each enzyme class. The ANNs model appeared best as it was seen to be relatively stable in performance through the training, cross-validation, and test sets of the data than the MLR model. For the prediction of Km, the most influential descriptors were partition coefficients and functional groups or fragments for compounds metabolised by ADH, ALDH, and FMO. Size, shape, symmetry, and atom distribution are those properties that mostly influenced the prediction of Vmax. This study is valuable in predicting Km and Vmax and for understanding the principles behind biotransformation by the liver enzymes; which in turn can be useful in taking proactive and remedial actions on issues regarding industrial activities affecting environmental wellbeing. It also finds relevance when guidance is needed for selecting an appropriate analytical model for a given dataset.

Keywords

Machine Learning, Supervised Learning, Artificial Neural Network, Multiple Linear Regression, Quantitative Structure-Activity Relationships, Xenobiotic, Michaelis-Menten Constant.

References

[1] Alessandra Pirovano, Stefan Brandmaier, Mark A. J. Huijbregts, Ad M. J. Ragas, Karin Veltman and A. Jan Hendriks (2015). The utilisation of structural descriptors to predict metabolic constants of xenobiotics in mammals. Environmental Toxicology and Pharmacology. 39: 247-258.

[2] Albert Chern (2015). “Introduction to MATLAB”. ACM11 Spring 2015, California Institute of Technology.

[3] Alexandre Varnek and Igor Baskin (2011). Machine Learning Methods for Property Prediction in Chemoinformatics: Quo Vadis?. Journal of Chemical Information and Modelling, dx.doi.org/10.1021/ci200409x.

[4] Ammar Abdo, Beining Chen, Christoph Mueller, Naomie Salim, and Peter Willett. (2010). Ligand-based virtual screening using Bayesian networks. J. Chem. Inf. Model. 50 (6) 1012–1020.

[5] Andrea Mauri, Viviana Consonni, Manuela Pavan, and Roberto Todeschini (2006). Dragon software: an easy approach to molecular descriptor calculations. MATCH Commun. Math. Comput. Chem. 56: 237-248, ISSN 0340 – 6253.

[6] Andreas Karoly Gombert and Jens Nielsen (2000). Mathematical modelling of metabolism. Current Opinion in Biotechnology, 11: 180–186.

[7] Andrew Ng. (2018). Coursera. Stanford Online Machine Learning Lecture.

[8] Antonio Lavecchia (2015). Machine-learning approaches in drug discovery: methods and applications. Drug Discovery Today, 20 (3) 318 – 331.

[9] Bailey J. E. (1998). Mathematical modelling and analysis in biochemical engineering: past accomplishments and future opportunities. Biotechnology Prog, 14: 8-20.

[10] Balaz, S. (2009). Modelling kinetics of subcellular disposition of chemicals. Chem. Rev. 109: 1793–1899.

[11] BioFoundations (2018). The Detoxification and Biotransformation System in the Human Body. https://biofoundations.org/the-detoxification-and-biotransformation-system-in-the-human-body/. Extracted on 29th March 2018.

[12] BRENDA: The Comprehensive Enzyme information system. https://www.brenda enzymes.org/.

[13] Byvatov, E. (2003). Comparison of support vector machine and artificial neural network systems for drug/nondrug classification. J. Chem. Inf. Comput. Sci. 43: 1882–1889.

[14] Chemical Computing Group. https://www.chemcomp.com/journal/descr.htm. Extracted on the 26th of January 2019.

[15] Cheng, T. et al. (2011). Binary classification of aqueous solubility using support vector machines with reduction and recombination feature selection. J. Chem. Inf. Model. 51: 229–236.

[16] Cherkasov, A., Muratov, E.N., Fourches, D., Varnek, A., Baskin, I.I.,Cronin, M., Dearden, J., Gramatica, P., Martin, Y.C., Todeschini,R., Consonni, V., Kuz’min, V.E., Cramer, R., Benigni, R., Yang,C., Rathman, J., Terfloth, L., Gasteiger, J., Richard, A., Tropsha,A. (2013). QSAR modelling: where have you been? Where are you going to?. J. Med. Chem. 57: 4977–5010.

[17] Consonni, V., Todeschini, R. (2010). Molecular descriptors. Recent Advances in QSAR Studies. Springer, Dordrecht, the Netherlands, pp. 29–102.

[18] David A. Winkler and Frank R. Burden (2000). “Robust QSAR Models from Novel Descriptors and Bayesian Regularised Neural Networks”. Molecular Simulation. 24: 4-6, 243-258, DOI: 10.1080/08927020008022374.

[19] Deconinck, E. et al. (2006). Classification tree models for the prediction of blood– brain barrier passage of drugs. Journal of Chem. Inf. Model. 46: 1410–1419.

[20] Dmitrij Martynenko (2015). “Big Data Analytics with MATLAB”. http://www.mathworks.com/discovery/matlab-mapreduce-hadoop.html. Extracted on 29th March 2018.

[21] Emre Karakoc, S. Cenk Sahinalp, and Artem Cherkasov (2006). Comparative QSAR – and Fragments Distribution Analysis of Drugs, Drug-likes, Metabolic Substances, and Antimicrobial Compounds. J. Chem. Inf. Model. 46: 2167-2182.

[22] Fogel, G.B. (2008). Computational intelligence approaches for pattern discovery in biological systems. Brief Bioinform. 9: 307–316.

[23] Foody, G.M. and Mathur, A. (2006). The use of small training sets containing mixed pixels for accurate hard image classification: training on mixed spectral responses for classification by SVM. Remote Sens. Environ. 103: 179–189.

[24] Frank R. Burden (1999). Robust QSAR Models Using Bayesian Regularized Neural Networks. J. Med. Chem., 42: 3183-3187.

[25] Frank, E. et al. (2000). Technical note: naive Bayes for regression. Mach. Learn. 41: 5–25

[26] Garrett, R., Grisham, C. M., (2010). Biochemistry, fourth ed. Brooks/Cole, Cengage Learning, Boston, MA, USA.

[27] GeorgeW. Bassel, Enrico Glaab, Julietta Marquez, Michael J. Holdsworth, and Jaume Bacardit (2011). Functional Network Construction in Arabidopsis Using Rule-Based Machine Learning on Large-Scale Data Sets. Large-Scale Biology Article, 23: 3101–3116.

[28] Gershenfeld N. A. (1999). The Nature of Mathematical Modelling. Cambridge: Cambridge University Press.

[29] Gianpaolo Bravi and James H. Wikel (2000). Application of MS‐WHIM Descriptors: 1. Introduction of New Molecular Surface Properties and 2. Prediction of Binding Affinity Data. Quant. Struct. Act. Relat., 19. https://doi.org/10.1002/(SICI)1521 3838(200002)19:1<29::AID-QSAR29>3.0.CO;2-P.

[30] Gleeson, M. P. et al. (2006). In silico human and rat Vss quantitative structure– activity relationship models. J. Med. Chem. 49: 1953–1963.

[31] Gregory Sliwoski, Jeffrey Mendenhall, and Jens Meiler (2015). Autocorrelation descriptor improvements for QSAR: 2DA_Sign and 3DA_Sign. J. Comput. Aided Mol Des. DOI 10.1007/s10822-015-9893-9.

[32] Haiping Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos (2011). A survey of multilinear subspace learning for tensor data. Pattern Recognition. 44: 1540–1551.

[33] Hansch, C., Mekapati, S.B., Kurup, A., Verma, R.P., (2004). QSAR of cytochrome P450. Drug. Metab. Rev. 36: 105–156.

[34] Haykin, S. S. (1999). Neural Networks: A Comprehensive Foundation. Prentice Hall.

[35] Ho, T. K. (1998). The random subspace method for constructing decision forests. ITPAM 20: 832–844.

[36] Hou, T. et al. (2007). ADME evaluation in drug discovery. 8. The prediction of human intestinal absorption by a support vector machine. Journal of Chem. Inf. Model. 47: 2408–2415.

[37] Iurii Sushko, Sergii Novotarskyi, Robert Ko¨rner, Anil Kumar Pandey, Matthias Rupp, Wolfram Teetz, Stefan Brandmaier, Ahmed Abdelaziz, Volodymyr V. Prokopenko, Vsevolod Y. Tanchuk, Roberto Todeschini, Alexandre Varnek, Gilles Marcou, Peter Ertl, Vladimir Potemkin, Maria Grishina, Johann Gasteiger, Christof Schwab, Igor I. Baskin, Vladimir A. Palyulin, Eugene V. Radchenko, William J. Welsh, Vladyslav Kholodovych, Dmitriy Chekmarev, Artem Cherkasov, Joao Aires-de-Sousa, Qing-You Zhang, Andreas Bender, Florian Nigsch, Luc Patiny, Antony Williams, Valery Tkachenko, Igor V. Tetko (2011). Online chemical modelling environment (OCHEM): web platform for data storage, model development and publishing of chemical information. J. Comput. Aided Mol Des. 25: 533–554, DOI 10.1007/s10822-011-9440-2.

[38] Jacek Kujawski, Marek K. Bernard, Anna Janusz, and Weronika Kuzma (2011). Prediction of log P: ALOGPS Application in Medicinal Chemistry Education. J. Chem. Educ. 2012, 89, 64–67. dx.doi.org/10.1021/ed100444h.

[39] Johannes Kirchmair, Andreas H. Göller, Dieter Lang, Jens Kunze, Bernard Testa, Ian D. Wilson, Robert C. Glen and Gisbert Schneider (2015). Predicting drug metabolism: experiment and/or computation?. PERSPECTIVES. 14: 389-404.

[40] Jun Zhang, Zhi-hui Zhan, Ying Lin, Ni Chen, Yue-jiao Gong, Jing-hui Zhong, Henry S.H. Chung, Yun Li, Yu-hui Shi (2011). Evolutionary Computation Meets Machine Learning: A Survey. IEEE Computational Intelligence Magazine, pp 68-75.

[41] Kathleen M. Knights, Andrew Rowland, and John O. Miners (2013). Renal drug metabolism in humans: The potential for drug-endobiotic interactions. British Journal of Clinical Pharmacology. 76 (4) 587-602.

[42] Kauffman, G. W. and Jurs, P.C. (2001). QSAR and k-nearest neighbour classification analysis of selective cyclooxygenase-2 inhibitors using topologically-based numerical descriptors. Journal of Chem. Inf. Comp. Sci. 41: 1553–1560.

[43] Krueger, S.K., Williams, D.E., (2005). Mammalian flavin-containing monooxygenases: structure/function, genetic polymorphisms and role in drug metabolism. Pharmacol. Ther. 106: 357–387.

[44] Lamanna, C. et al. (2008). Straightforward recursive partitioning model for discarding insoluble compounds in the drug discovery process. Journal of Med. Chem. 51, 2891–2897

[45] Lash, Lawrence H. (1994). “Role of Renal Metabolism in Risk”. Environmental Health Perspectives. 102 (11) 75-79.

[46] Lewis, D.F.V., (1999). Frontier orbitals in chemical and biological activity: quantitative relationships and mechanistic implication. Drug. Metab. Rev. 31: 755–816.

[47] List of molecular descriptors calculated by DRAGON. http://www.talete.mi.it/products/dragon_molecular_descriptor_list.pdf. Extracted on 30th January 2019.

[48] Lowe, R. et al. (2012). Predicting the mechanism of phospholipidosis. Journal of Cheminformatics 4: 2.

[49] Marco Chiarandini. “Machine Learning: Linear Regression and Neural Networks”. Introduction to Computer Science. Department of Mathematics & Computer Science University of Southern Denmark.

[50] Margot Gerritsen (2006). “A brief introduction to MATLAB”. Linear Algebra with Application to Engineering Computations, Autumn 2006 Handout 3.

[51] MathWorks (2016). “Introducing Machine Learning”. mathwork.com/trademarks. Extracted on 2nd April 2018.

[52] Mayer-Schönberger, V., and Cukier, K. (2014). Big data: A revolution that will transform how. American Journal of Epidemiology. 179 (9) 1143–1144.

[53] Mente, S. R. et al. (2005). A recursive-partitioning model for blood–brain barrier permeation. J. Comput. Aided Mol. Des. 19: 465–481:

[54] Nielsen J., Jørgensen H. S. (1996). A kinetic model for the penicillin biosynthetic pathway in Penicillium chrysogenum. Control Eng Practice, 4:765-771.

[55] Nigsch, F. et al. (2006). Melting point prediction employing k-nearest neighbour algorithms and genetic parameter optimization. J. Chem. Inf. Model. 46: 2412–2422.

[56] Oleg Devinyak, Dmytro Havrylyuk, and Roman Lesyk (2014). 3D-MoRSE Descriptors Explained. Journal of Molecular Graphics and Modelling. DOI: 10.1016/j.jmgm.2014.10.006.

[57] Patel, J. and Chaudhari, C. (2005). Introduction to the artificial neural networks and their applications in QSAR studies. ALTEX. 22: 271.

[58] Pirovano, A., Huijbregts, M.A.J., Ragas, A.M.J., Veltman, K.,Hendriks, A.J. (2014). Mechanistically-based QSARs to describe metabolic constants in mammals. ATLA. 42: 59–69.

[59] Pissara P. N., Nielsen J., Bazin M. J. (1996). Pathway kinetics and metabolic control analysis of a high-yielding strain of Penicillium chrysogenum during fed batch cultivations. Biotechnology Bioeng, 51:168-176.

[60] Rizzi M, Baltes M, Theobald U, Reuss M (1997). In vivo analysis of metabolic dynamics in Saccharomyces cerevisiae: II. mathematical model. Biotechnol Bioeng, 55:592-608.

[61] S. Agatonovic-Kustrin and R. Beresford (2000). Basic concepts of artificial neural network (ANN) modelling and its application in pharmaceutical research. Journal of Pharmaceutical and Biomedical Analysis. 22 (5) 717-727.

[62] Sakiyama, Y. et al. (2008). Predicting human liver microsomal stability with machine learning techniques. J. Mol. Graph. Model. 26: 907–915.

[63] Scheer, M., Grote, A., Chang, A., Schomburg, I., Munaretto, C.,Rother, M., Söhngen, C., Stelzer, M., Thiele, J., Schomburg, D. (2011). “BRENDA, the enzyme information system”. Nucleic Acids Res. 39: D670–D676.

[64] Schilling C. H., Edwards J. S., Palsson B. O. (1999). Toward metabolic phenomics: analysis of genomic data using flux balances. Biotechnol. Prog. 15: 288-295.

[65] Shashi K. Ramaiah and Atrayee Banerjee (2015). “Liver Toxicity of Chemical Warfare Agents”: Handbook of Toxicology of Chemical Warfare. ScienceDirect. Pp 615-626.

[66] Theilgaard H., Nielsen J. (1999). Metabolic control analysis of the penicillin biosynthetic pathway: the influence of the LLD-ACV: bis ACV ratio on the flux control. Anton Leeuw Int J G, 75: 145-154.

[67] Tiago M. Fragoso and Francisco Louzada Neto (2017). Bayesian model averaging: A systematic review and conceptual classification. International Statistical Review. 0(0)1–28. doi:10.1111/insr.12243.

[68] Tropsha, Alexander (2010). Best Practices for QSAR Model Development, Validation, and Exploitation. Molecular Informatics. 29: 6-7: 476–488.

[69] Uthayasankar Sivarajah, Muhammad Mustafa Kamal, Zahir Irani, Vishanth Weerakkody (2016). Critical analysis of Big Data challenges and analytical methods. Journal of Business Research. http://dx.doi.org/10.1016/j.jbusres.2016.08.001.

[70] Vapnik, V. N. (1998). Statistical Learning Theory. Wiley.

[71] Vapnik, V. N. (2000). The Nature of Statistical Learning Theory. Springer

[72] Vasiliou, V., Pappa, A., Petersen, D.R., (2000). Role of aldehyde dehydrogenases in endogenous and xenobiotic metabolism. Chem. Biol. Interact. 129: 1–19.

[73] Viviana Consonni, Roberto Todeschini, Manuela Pavan, and Paola Gramatica (2002). Structure/Response Correlations and Similarity/Diversity Analysis by GETAWAY Descriptors. 2. Application of the Novel 3D Molecular Descriptors to QSAR/QSPR Studies. J. Chem. Inf. Comput. Sci. 42: 693-705.

[74] Von Korff, M. and Sander, T. (2006). Toxicity-indicating structural patterns. J. Chem. Inf. Model. 46: 536–544.

[75] Waller, C.L., Evans, M.V., and McKinney, J.D. (1996). Modelling the cytochrome P450-mediated metabolism of chlorinated volatile organic compounds. Drug Metab. Dispos.24: 203–210.

[76] Wasserman, L. (2000). Bayesian model selection and model averaging. Journal of Mathematical Psychology. 44: 92–107.

[77] Wilbert B. Copeland, Bryan A. Bartley, Deepak Chandran, Michal Galdzicki, Kyung H. Kim, Sean C. Sleight, Costas D. Maranas, Herbert M. Sauro (2012). Computational tools for metabolic engineering. Metabolic Engineering. 14: 270–280.

[78] Willett, P. et al. (2007). Prediction of ion channel activity using binary kernel discrimination. J. Chem. Inf. Model. 47: 1961–1966.

[79] Yousefinejad S. and Hemmateenejad B. (2015). "Chemometrics tools in QSAR/QSPR studies: A historical perspective". Chemometric and Intelligent Laboratory Systems. Part B, 149: 177–204.

[80] Zvinavashe, E., Murk, A.J., Rietjens, I.M.C.M., (2008). “Promises and pitfalls of quantitative structure–activity relationship approaches for predicting metabolism and toxicity”. Chem. Res. Toxicol. 21, 2229–2236.

How to cite this paper

Sunday Tarekakpo Odobai, Nazifi Lawal Bashir "Data Analytics ? Computer Modelling of Metabolic Rates" Iconic Research And Engineering Journals Volume 5 Issue 8 2022 Page 116-132
Sunday Tarekakpo Odobai, Nazifi Lawal Bashir "Data Analytics ? Computer Modelling of Metabolic Rates" Iconic Research And Engineering Journals, vol. 5, no. 8, Feb. 2022
Sunday Tarekakpo Odobai, Nazifi Lawal Bashir (2022). Data Analytics ? Computer Modelling of Metabolic Rates. Iconic Research And Engineering Journals, 5(8).
Sunday Tarekakpo Odobai, Nazifi Lawal Bashir "Data Analytics ? Computer Modelling of Metabolic Rates" Iconic Research And Engineering Journals, vol. 5, no. 8, Feb. 2022.
@article{1703191,
      author = {Sunday Tarekakpo Odobai, Nazifi Lawal Bashir},
      title = {Data Analytics ? Computer Modelling of Metabolic Rates},
      journal = {Iconic Research And Engineering Journals},
      year = {2022},
      volume = {5},
      number = {8},
      pages = {116-132},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1703191.pdf},
      abstract = {Artificial Neural Networks (ANNs) and Multiple Linear Regression (MLR) based Quantitative Structure-Activity Relationships (QSARs) models were developed to predict enzymatic activities, that is, the Michaelis-Menten constant (Km) and the maximum reaction rate (Vmax) for reactions involving the biotransformation of xenobiotics, catalysed by three classes of enzymes present in the mammalian livers. The enzymes we have studied here are alcohol dehydrogenase (ADH), aldehyde dehydrogenase (ALDH), and Flavin-containing monooxygenase (FMO). Data for enzymatic constants were collected from the literature and the computation of potential predictors was done for all xenobiotics to include for hundreds of molecular descriptors. The best predictor variables were selected (maximum of seven and a minimum of two descriptors) using the Microsoft excel correlation function for each enzyme class. Each dataset was divided into three sets, the divisions were training, cross-validation, and test sets in the ratio of 70%, 15%, and 15% respectively for both the ANNs and the MLR models to build the QSARs. The MATLAB programming language was employed to implement the writing and running of the learning algorithms. The predictive strengths of the models were assessed through the correlation of their predictions relative to the target outcomes for the three divisions and the mean square errors were computed, after fitting the resulting models with the entire dataset for each enzyme class. The ANNs model appeared best as it was seen to be relatively stable in performance through the training, cross-validation, and test sets of the data than the MLR model. For the prediction of Km, the most influential descriptors were partition coefficients and functional groups or fragments for compounds metabolised by ADH, ALDH, and FMO. Size, shape, symmetry, and atom distribution are those properties that mostly influenced the prediction of Vmax. This study is valuable in predicting Km and Vmax and for understanding the principles behind biotransformation by the liver enzymes; which in turn can be useful in taking proactive and remedial actions on issues regarding industrial activities affecting environmental wellbeing. It also finds relevance when guidance is needed for selecting an appropriate analytical model for a given dataset.},
      keywords = {Machine Learning, Supervised Learning, Artificial Neural Network, Multiple Linear Regression, Quantitative Structure-Activity Relationships, Xenobiotic, Michaelis-Menten Constant. },
      month = {February},
  }