The recent worldwide spreading of pneumonia-causing virus, such as Coronavirus, COVID-19, and H1N1, has been endangering the life of human beings all around the world. In order to really understand the biological process within a cell level and provide useful clues to develop antiviral drugs, information of virus protein subcellular localization is vitally important. In view of this, a CNN based virus protein subcellular localization predictor called “pLoc_Deep-mVirus” was developed. The predictor is particularly useful in dealing with the multi-sites systems in which some proteins may simultaneously occur in two or more different organelles that are the current focus of pharmaceutical industry. The global absolute true rate achieved by the new predictor is over 97% and its local accuracy is over 98%. Both are transcending other existing state-of-the-art predictors significantly. It has not escaped our notice that the deep-learning treatment can be used to deal with many other biological systems as well. To maximize the convenience for most experimental scientists, a user-friendly web-server for the new predictor has been established at http://www.jci-bioinfo.cn/pLoc_Deep-mVirus/ .
Ehrlich, J.S., Hansen, M.D. and Nelson, W.J. (2002) Spatio-Temporal Regulation of Rac1 Localization and Lamellipodia Dynamics during Epithelial Cell-Cell Adhesion. Developmental Cell, 3, 259-270. https://doi.org/10.1016/S1534-5807(02)00216-2
Glory, E. and Murphy, R.F. (2007) Automated Subcellular Location Determination and High-Throughput Microscopy. Developmental Cell, 12, 7-16. https://doi.org/10.1016/j.devcel.2006.12.007
Chou, K.C. (2015) Impacts of Bioinformatics to Medicinal Chemistry. Medicinal Chemistry, 11, 218-234. https://doi.org/10.2174/1573406411666141229162834
Xiao, X., Cheng, X., Chen, G., Mao, Q. and Chou, K.C. (2019) pLoc_bal-mVirus: Predict Subcellular Localization of Multi-Label Virus Proteins by Chou’s General PseAAC and IHTS Treatment to Balance Training Dataset. Medicinal Chemistry, 15, 496-509. https://doi.org/10.2174/1573406415666181217114710
Nakai, K. and Kanehisa, M. (1992) A Knowledge Base for Predicting Protein Localization Sites in Eukaryotic Cells. Genomics, 14, 897-911. https://doi.org/10.1016/S0888-7543(05)80111-9
Cedano, J., Aloy, P., Perez-Pons, J.A. and Querol, E. (1997) Relation between Amino Acid Composition and Cellular Location of Proteins. Journal of Molecular Biology, 266, 594-600. https://doi.org/10.1006/jmbi.1996.0804
Reinhardt, A. and Hubbard, T. (1998) Using Neural Networks for Prediction of the Subcellular Location of Proteins. Nucleic Acids Research, 26, 2230-2236. https://doi.org/10.1093/nar/26.9.2230
Chou, K.C. and Shen, H.B. (2007) Recent Progresses in Protein Subcellular Location Prediction. Analytical Biochemistry, 370, 1-16. https://doi.org/10.1016/j.ab.2007.07.006
Chou, K.C., Wu, Z.C. and Xiao, X. (2011) iLoc-Euk: A Multi-Label Classifier for Predicting the Subcellular Localization of Singleplex and Multiplex Eukaryotic Proteins. PLoS ONE, 6, e18258. https://doi.org/10.1371/journal.pone.0018258
Mandal, M., Mukhopadhyay, A. and Maulik, U. (2015) Prediction of Protein Subcellular Localization by Incorporating Multiobjective PSO-Based Feature Subset Selection into the General form of Chou’s PseAAC. Medical & Biological Engineering & Computing, 53, 331-344. https://doi.org/10.1007/s11517-014-1238-7
Maxwell, A., Li, R., Yang, B., Weng, H., Ou, A., Hong, H., Zhou, Z., Gong, P. and Zhang, C. (2017) Deep Learning Architectures for Multi-Label Classification of Intelligent Health Risk Prediction. BMC Bioinformatics, 18, 523. https://doi.org/10.1186/s12859-017-1898-z
Khan, S., Khan, M., Iqbal, N., Hussain, T., Khan, S.A. and Chou, K.C. (2019) A Two-Level Computation Model Based on Deep Learning Algorithm for Identification of piRNA and Their Functions via Chou’s 5-Steps Rule. International Journal of Peptide Research and Therapeutics, 26, 795-809. https://doi.org/10.1007/s10989-019-09887-3
Khan, Z.U., Ali, F., Khan, I.A., Hussain, Y. and Pi, D. (2019) iRSpot-SPI: Deep Learning-Based Recombination Spots Prediction by Incorporating Secondary Sequence Information Coupled with Physio-Chemical Properties via Chou’s 5-Step Rule and Pseudo Components. Chemometrics and Intelligent Laboratory Systems (CHEMOLAB), 189, 169-180. https://doi.org/10.1016/j.chemolab.2019.05.003
Nazari, I., Tahir, M., Tayari, H. and Chong, K.T. (2019) iN6-Methyl (5-Step): Identifying RNA N6-Methyladenosine Sites Using Deep Learning Mode via Chou’s 5-Step Rules and Chou’s General PseKNC. Chemometrics and Intelligent Laboratory Systems (CHEMOLAB), 189, 169-180. https://doi.org/10.1016/j.chemolab.2019.103811
Hussain, W., Khan, Y.D., Rasool, N., Khan, S.A. and Chou, K.C. (2019) SPrenylC-PseAAC: A Sequence-Based Model Developed via Chou’s 5-Steps Rule and General PseAAC for Identifying S-Prenylation Sites in Proteins. Journal of Theoretical Biology, 468, 195-203. https://doi.org/10.1016/j.jtbi.2019.02.007
Charoenkwan, P., Schaduangrat, N., Nantasenamat, C., Piacham, T. and Shoombuatong, W. (2020) iQSP: A Sequence-Based Tool for the Prediction and Analysis of Quorum Sensing Peptides via Chou’s 5-Steps Rule and Informative Physicochemical Properties. International Journal of Molecular Sciences, 21, 75. https://doi.org/10.3390/ijms21010075
Chou, K.C. (2011) Some Remarks on Protein Attribute Prediction and Pseudo Amino Acid Composition (50th Anniversary Year Review, 5-Steps Rule). Journal of Theoretical Biology, 273, 236-247. https://doi.org/10.1016/j.jtbi.2010.12.024
Shen, H.B. and Chou, K.C. (2010) Virus-mPLoc: A Fusion Classifier for Viral Protein Subcellular Location Prediction by Incorporating Multiple Sites. Journal of Biomolecular Structure and Dynamics (JBSD), 28, 175-186. https://doi.org/10.1080/07391102.2010.10507351
Chou, K.C. (2001)d Prediction of Protein Cellular Attributes Using Pseudo Amino Acid Composition. PROTEINS: Structure, Function, and Genetics, 43, 246-255. (Erratum: ibid., 2001, Vol. 44, 60) https://doi.org/10.1002/prot.1035
Chou, K.C. (2005) Using Amphiphilic Pseudo Amino Acid Composition to Predict Enzyme Subfamily Classes. Bioinformatics, 21, 10-19. https://doi.org/10.1093/bioinformatics/bth466
Zhou, X.B., Chen, C., Li, Z.C. and Zou, X.Y. (2007) Using Chou’s Amphiphilic Pseudo Amino Acid Composition and Support Vector Machine for Prediction of Enzyme Subfamily Classes. Journal of Theoretical Biology, 248, 546-551. https://doi.org/10.1016/j.jtbi.2007.06.001
Zhang, S.W., Chen, W., Yang, F. and Pan, Q. (2008) Using Chou’s Pseudo Amino Acid Composition to Predict Protein Quaternary Structure: A Sequence-Segmented PseAAC Approach. Amino Acids, 35, 591-598. https://doi.org/10.1007/s00726-008-0086-x
Qiu, J.D., Huang, J.H., Liang, R.P. and Lu, X.Q. (2009) Prediction of G-Protein-Coupled Receptor Classes Based on the Concept of Chou’s Pseudo Amino Acid Composition: An Approach from Discrete Wavelet Transform. Analytical Biochemistry, 390, 68-73. https://doi.org/10.1016/j.ab.2009.04.009
Mohabatkar, H. (2010) Prediction of Cyclin Proteins Using Chou’s Pseudo Amino Acid Composition. Protein & Peptide Letters, 17, 1207-1214. https://doi.org/10.2174/092986610792231564
Qiu, J.D., Suo, S.B., Sun, X.Y., Shi, S.P. and Liang, R.P. (2011) OligoPred: A Web-Server for Predicting Homo-Oligomeric Proteins by Incorporating Discrete Wavelet Transform into Chou’s Pseudo Amino Acid Composition. Journal of Molecular Graphics & Modelling, 30, 129-134. https://doi.org/10.1016/j.jmgm.2011.06.014
Nanni, L., Lumini, A., Gupta, D. and Garg, A. (2012) Identifying Bacterial Virulent Proteins by Fusing a Set of Classifiers Based on Variants of Chou’s Pseudo Amino Acid Composition and on Evolutionary Information. IEEE-ACM Transaction on Computational Biolology and Bioinformatics, 9, 467-475. https://doi.org/10.1109/TCBB.2011.117
Khosravian, M., Faramarzi, F.K., Beigi, M.M., Behbahani, M. and Mohabatkar, H. (2013) Predicting Antibacterial Peptides by the Concept of Chou’s Pseudo amino Acid Composition and Machine Learning Methods. Protein & Peptide Letters, 20, 180-186. https://doi.org/10.2174/092986613804725307
Kumar, R., Srivastava, A., Kumari, B. and Kumar, M. (2015) Prediction of Beta-Lactamase and Its Class by Chou’s Pseudo Amino Acid Composition and Support Vector Machine. Journal of Theoretical Biology, 365, 96-103. https://doi.org/10.1016/j.jtbi.2014.10.008
Mei, J., Fu, Y. and Zhao, J. (2018) Analysis and Prediction of Ion Channel Inhibitors by Using Feature Selection and Chou’s General Pseudo Amino Acid Composition. Journal of Theoretical Biology, 456, 41-48. https://doi.org/10.1016/j.jtbi.2018.07.040
Zhang, S., Yang, K., Lei, Y. and Song, K. (2019) iRSpot-DTS: Predict Recombination Spots by Incorporating the Dinucleotide-Based Spare-Cross Covariance Information into Chou’s Pseudo Components. Genomics, 111, 1760-1770. https://doi.org/10.1016/j.ygeno.2018.11.031
Akbar, S., Rahman, A.U. and Hayat, M. (2020) cACP: Classifying Anticancer Peptides Using Discriminative Intelligent Model via Chou’s 5-Step Rules and General Pseudo Components. Chemometrics and Intelligent Laboratory (CHEMOLAB), 196, Article ID: 103912. https://doi.org/10.1016/j.chemolab.2019.103912
Chou, K.C. (2017) An Unprecedented Revolution in Medicinal Chemistry Driven by the Progress of Biological Science. Current Topics in Medicinal Chemistry, 17, 2337-2358. https://doi.org/10.2174/1568026617666170414145508
Du, P., Wang, X., Xu, C. and Gao, Y. (2012) PseAAC-Builder: A Cross-Platform Stand-Alone Program for Generating Various Special Chou’s Pseudo Amino Acid Compositions. Analytical Biochemistry, 425, 117-119. https://doi.org/10.1016/j.ab.2012.03.015
Cao, D.S., Xu, Q.S. and Liang, Y.Z. (2013) Propy: A Tool to Generate Various Modes of Chou’s PseAAC. Bioinformatics, 29, 960-962. https://doi.org/10.1093/bioinformatics/btt072
Du, P., Gu, S. and Jiao, Y. (2014) PseAAC-General: Fast Building Various Modes of General form of Chou’s Pseudo Amino Acid Composition for Large-Scale Protein Datasets. International Journal of Molecular Sciences, 15, 3495-3506. https://doi.org/10.3390/ijms15033495
Chou, K.C. (2009) Pseudo Amino Acid Composition and Its Applications in Bioinformatics, Proteomics and System Biology. Current Proteomics, 6, 262-274. https://doi.org/10.2174/157016409789973707
Chen, W., Lei, T.Y., Jin, D.C., Lin, H. and Chou, K.C. (2014) PseKNC: A Flexible Web-Server for Generating Pseudo K-Tuple Nucleotide Composition. Analytical Biochemistry, 456, 53-60. https://doi.org/10.1016/j.ab.2014.04.001
Chen, W., Lin, H. and Chou, K.C. (2015) Pseudo Nucleotide Composition or PseKNC: An Effective Formulation for Analyzing Genomic Sequences. Molecular BioSystems, 11, 2620-2634. https://doi.org/10.1039/C5MB00155B
Liu, B., Yang, F. and Chou, K.C. (2017) 2L-piRNA: A Two-Layer Ensemble Classifier for Identifying Piwi-Interacting RNAs and Their Function. Molecular Therapy—Nucleic Acids, 7, 267-277. https://doi.org/10.1016/j.omtn.2017.04.008
Glorot, X., Bordes, A. and Bengio, Y. (2011) Deep Sparse Rectifier Neural Networks. 14th International Conference on Artificial Intelligence and Statistics, Ft. Lauderdale, 11-13 April 2011, 315-323.
Chou, K.C. (2019) Two Kinds of Metrics for Computational Biology. Genomics. https://www.sciencedirect.com/science/article/pii/S0888754319304604?via%3Dihub
Chou, K.C. (2013) Some Remarks on Predicting Multi-Label Attributes in Molecular Biosystems. Molecular Biosystems, 9, 1092-1100. https://doi.org/10.1039/c3mb25555g
Song, J., Wang, Y., Li, F., Akutsu, T., Rawlings, N.D., Webb, G.I. and Chou, K.C. (2018) iProt-Sub: A Comprehensive Package for Accurately Mapping and Predicting Protease-Specific Substrates and Cleavage Sites. Brief in Bioinform, 20, 638-658. https://doi.org/10.1093/bib/bby028
Zhang, M., Li, F., Marquez-Lago, T.T., Leier, A., Fan, C., Kwoh, C.K., Chou, K.C., Song, J. and Jia, C. (2019) MULTiPly: A Novel Multi-Layer Predictor for Discovering General and Specific Types of Promoters. Bioinformatics, 35, 2957-2965. https://doi.org/10.1093/bioinformatics/btz016
Shen, H.B. and Chou, K.C. (2007) Hum-mPLoc: An Ensemble Classifier for Large-Scale Human Protein Subcellular Location Prediction by Incorporating Samples with Multiple Sites. Biochemical and Biophysical Research Communications (BBRC), 355, 1006-1011. https://doi.org/10.1016/j.bbrc.2007.02.071
Chou, K.C. and Shen, H.B. (2008) Cell-PLoc: A Package of Web Servers for Predicting Subcellular Localization of Proteins in Various Organisms. Nature Protocols, 3, 153-162. https://doi.org/10.1038/nprot.2007.494
Shen, H.B. and Chou, K.C. (2009) A Top-Down Approach to Enhance the Power of Predicting Human Protein Subcellular Localization: Hum-mPLoc 2.0. Analytical Biochemistry, 394, 269-274. https://doi.org/10.1016/j.ab.2009.07.046
Chou, K.C. and Shen, H.B. (2010) Cell-PLoc 2.0: An Improved Package of Web-Servers for Predicting Subcellular Localization of Proteins in Various Organisms. Natural Science, 2, 1090-1103. https://doi.org/10.4236/ns.2010.210136
Chou, K.C., Wu, Z.C. and Xiao, X. (2012) iLoc-Hum: Using Accumulation-Label Scale to Predict Subcellular Locations of Human Proteins with Both Single and Multiple Sites. Molecular Biosystems, 8, 629-641. https://doi.org/10.1039/C1MB05420A
Cheng, X., Xiao, X. and Chou, K.C. (2018) pLoc-mHum: Predict Subcellular Localization of Multi-Location Human Proteins via General PseAAC to Winnow out the Crucial GO Information. Bioinformatics, 34, 1448-1456. https://doi.org/10.1093/bioinformatics/btx711
Cao, J.Z., Liu, W.Q. and Gu, H. (2012) Predicting Viral Protein Subcellular Localization with Chou’s Pseudo Amino Acid Composition and Imbalance-Weighted Multi-Label K-Nearest Neighbor Algorithm. Protein and Peptide Letters, 19, 1163-1169. https://doi.org/10.2174/092986612803216999
He, J., Gu, H. and Liu, W. (2012) Imbalanced Multi-Modal Multi-Label Learning for Subcellular Localization Prediction of Human Proteins with Both Single and Multiple Sites. PLoS ONE, 7, e37155. https://doi.org/10.1371/journal.pone.0037155
Li, L.Q., Zhang, Y., Zou, L.Y., Zhou, Y. and Zheng, X.Q. (2012) Prediction of Protein Subcellular Multi-Localization Based on the General form of Chou’s Pseudo Amino Acid Composition. Protein & Peptide Letters, 19, 375-387. https://doi.org/10.2174/092986612799789369
Mei, S. (2012) Predicting Plant Protein Subcellular Multi-Localization by Chou’s PseAAC Formulation Based Multi-Label Homolog Knowledge Transfer Learning. Journal of Theoretical Biology, 310, 80-87. https://doi.org/10.1016/j.jtbi.2012.06.028
Wang, X. and Li, G.Z. (2012) A Multi-Label Predictor for Identifying the Subcellular Locations of Singleplex and Multiplex Eukaryotic Proteins. PLoS ONE, 7, e36317. https://doi.org/10.1371/journal.pone.0036317
Huang, C. and Yuan, J. (2013) Using Radial Basis Function on the General Form of Chou’s Pseudo Amino Acid Composition and PSSM to Predict Subcellular Locations of Proteins with Both Single and Multiple Sites. Biosystems, 113, 50-57. https://doi.org/10.1016/j.biosystems.2013.04.005
Pacharawongsakda, E. and Theeramunkong, T. (2013) Predict Subcellular Locations of Singleplex and Multiplex Proteins by Semi-Supervised Learning and Dimension-Reducing General Mode of Chou’s PseAAC. IEEE Transactions on Nanobioscience, 12, 311-320. https://doi.org/10.1109/TNB.2013.2272014
Chen, W., Feng, P.M., Lin, H. and Chou, K.C. (2013) iRSpot-PseDNC: Identify Recombination Spots with Pseudo Dinucleotide Composition. Nucleic Acids Research, 41, e68. https://doi.org/10.1093/nar/gks1450
Chou, K.C. (2001) Using Subsite Coupling to Predict Signal Peptides. Protein Engineering, 14, 75-79. https://doi.org/10.1093/protein/14.2.75
Qiu, W.R., Xiao, X. and Chou, K.C. (2014) iRSpot-TNCPseAAC: Identify Recombination Spots with Trinucleotide Composition and Pseudo Amino Acid Components. International Journal of Molecular Sciences (IJMS), 15, 1746-1766. https://doi.org/10.3390/ijms15021746
Xu, R., Zhou, J., Liu, B., He, Y.A., Zou, Q., Wang, X. and Chou, K.C. (2015) Identification of DNA-Binding Proteins by Incorporating Evolutionary Information into Pseudo Amino Acid Composition via the Top-n-Gram Approach. Journal of Biomolecular Structure & Dynamics (JBSD), 33, 1720-1730. https://doi.org/10.1080/07391102.2014.968624
Jia, J., Zhang, L., Liu, Z., Xiao, X. and Chou, K.C. (2016) pSumo-CD: Predicting Sumoylation Sites in Proteins with Covariance Discriminant Algorithm by Incorporating Sequence-Coupled Effects into General PseAAC. Bioinformatics, 32, 3133-3141. https://doi.org/10.1093/bioinformatics/btw387
Liu, B., Wang, S., Long, R. and Chou, K.C. (2017) iRSpot-EL: Identify Recombination Spots with an Ensemble Learning Approach. Bioinformatics, 33, 35-41. https://doi.org/10.1093/bioinformatics/btw539
Chou, K.C. and Shen, H.B. (2009) Recent Advances in Developing Web-Servers for Predicting Protein Attributes. Natural Science, 1, 63-92. https://doi.org/10.4236/ns.2009.12011