Prediction of Protein Expression and Growth Rates by Supervised Machine Learning
- 1 School of Informatics, University of Edinburgh, Edinburgh, United Kingdom
Abstract
The DNA sequences of an organism play an important influence on its transcription and translation process, thus affecting its protein production and growth rate. Due to the com-plexity of DNA, it was extremely difficult to predict the macroscopic characteristics of or-ganisms. However, with the rapid development of machine learning in recent years, it be-comes possible to use powerful machine learning algorithms to process and analyze biolog-ical data. Based on the synthetic DNA sequences of a specific microbe, E. coli , I designed a process to predict its protein production and growth rate. By observing the properties of a data set constructed by previous work, I chose to use supervised learning regressors with encoded DNA sequences as input features to perform the predictions. After comparing different encoders and algorithms, I selected three encoders to encode the DNA sequences as inputs and trained seven different regressors to predict the outputs. The hy-per-parameters are optimized for three regressors which have the best potential prediction performance. Finally, I successfully predicted the protein production and growth rates, with the best R 2 score 0.55 and 0.77, respectively, by using encoders to catch the potential fea-tures from the DNA sequences.
- Tarca, A.L., Carey, V.J., Chen, W.X., Romero, R. and Draghici, S. (2007) Machine Learning and Its Applications to Biology. PLoS Computational Biology, 3, Article No. e116. https://doi.org/10.1371/journal.pcbi.0030116
- Sinden, R.R. (1994) DNA Structure and Function. Academic Press, Cambridge, 11-12.
- Henderson, J.F. and Paterson, A.R.P. (1973) Nucleotide Metabolism: An Introduction. Academic Press, Cambridge, 23-25.
- Stormo, G.D. and Zhao, Y. (2010) Determining the Specificity of Protein-DNA Interactions. Nature Reviews Genetics volume, 11, 751-760. https://doi.org/10.1038/nrg2845
- Riggs, P. (2021) What is mRNA? The Messenger Molecule That’s Been in Every Living Cell for Billions of Years Is the Key Ingredient in Some Covid-19 Vaccines. The Conversation.
- Guillaume, Cambray, Guimaraes, J.C. and Arkin, A.P. (2018) Evaluation of 244,000 Synthetic Sequences Reveals Design Principles to Optimize Translation in Escherichia Coli. Nature Biotechnology, 36, 1005-1015. https://doi.org/10.1038/nbt.4238
- Addgene (2017) Promoters. https://www.addgene.org/mol-bio-reference/promoters/.
- Ng, P. (2017) Dna2vec: Consistent Vector Representations of Variable-Length k-Mers. arXiv:1701.06279.
- Weisberg, S. (1973) Applied Linear Regression. John Wiley Sons, Inc., Hoboken, 19-33.
- Sharma, A. (2020) Decision Tree vs. Random Forest—Which Algorithm Should You Use? Analytics Vidhya.
- Oshiro, T.M., PerezJose′, P.S. and Baranauskas, A. (2012) How Many Trees in a Random Forest? International Workshop on Machine Learning and Data Mining in Pattern Recognition, Berlin, 13-20 July 2012, 154-168. https://doi.org/10.1007/978-3-642-31537-4_13
- Scikit Learn (2008) Neural Network Models (Supervised). https://scikit-learn.org/stable/modules/neural_networks_supervised.html
- Seiffert, U. (2001) Multiple Layer Perceptron Training Using Genetic Algorithms. European Symposium on Artificial Neural Networks, Bruges, 25-27 April 2001, 159-164.
- Peterson, L.E. (2009) K-Nearest Neighbor. Scholarpedia, 4, 1883.
- Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W. and Liu, T.-Y. (2017) Light-GBM: A Highly Efficient Gradient Boosting Decision Tree. Advances in Neural Information Processing Systems 30, Long Beach, 4-9 December 2017, 3148-3156.