Cancer Specific Non-Synonymous Single Nucleotide Polymorphism Prediction in the Context of Haplotype and Protein Interacting Sites — Oak Academic Publishing
Research ArticleOpen AccessGoogle Scholar indexed
Cancer Specific Non-Synonymous Single Nucleotide Polymorphism Prediction in the Context of Haplotype and Protein Interacting Sites
Computer and Information Sciences, University of Delaware, Newark, Delaware, USA
,
Computer and Information Sciences, University of Delaware, Newark, Delaware, USA
1 Computer and Information Sciences, University of Delaware, Newark, Delaware, USA
2 Computer and Information Sciences, University of Delaware, Newark, Delaware, USA
In this work, we study predicting the effect of non-synonymous SNPs on several cancers. We trained classifiers on both sequential and structural features extracted from the affected genes and assessed the predictions made by the trained classifiers using cross validation. Specifically, we investigated how the prediction performance can be improved by connecting SNPs in the context of haplotype and interacting sites of proteins encoded by affected genes. We found that accuracy was consistently enhanced by combining sequential and structural features, with increase ranging from a few percentage points up to more than 20 percentage points. The results for putting SNPs in the context of interacting sites were less consistent. Compared to individual SNPs, these that appear together in haplotype showed stronger correlation with one another and with the phenotype, and therefore led to significant improvement inprediction performance, with ROC score increased from 0.81 to 0.95. Although some similar effect has been expected for connecting SNPs to interacting sites in proteins, the performance actually got worse. This decrease in prediction accuracy may be caused by the small data set being used in the study, as many affected proteins in the study do not have known interacting sites.
Wu, J., Gan, M., and Jiang, R. (2011) Prioritisation of Candidate Single Amino Acid Polymorphisms Using One-Class Learning Machines. International Journal of Computational Biology and Drug Design, 4, 316–331. https://doi.org/10.1504/IJCBDD.2011.044446
David, A., Razali, R., Wass, M.N. and Sternberg, M.J. (2012) Protein-protein Interaction Sites are Hotspots for Disease-Associated Nonsynonymous SNPs. Hum Mutat 33, 359–363. https://doi.org/10.1002/humu.21656
Basit, N. and Wechsler, H. (2011) Prediction of Enzyme Mutant Activity Using Computa-tional Mutagenesis and Incremental Transduction. Advances in bioinformatics. Adv Bioinformatics, 2011, 958129. https://doi.org/10.1155/2011/958129
Lee, T.S. and York, D.M. (2010) Computational Mutagenesis Studies of Hammerhead Ribozyme Catalysis. J Am ChemSoc, 132, 13505–13518. https://doi.org/10.1021/ja105956u
Masso, M. and Vaisman, II. (2010) Knowledge-based Computational Mutagenesis for Pre-dicting the Disease Potential of Human Non-Synonymous Single Nucleotide Polymorphisms. J TheorBiol , 266, 560–568. https://doi.org/10.1016/j.jtbi.2010.07.026
Bradshaw, R.T., Patel, B.H., Tate, E.W., Leatherbarrow, R.J. and Gould, I.R. (2011) Comparing Experimental and Computational Alanine Scanning Techniques for Probing a Prototypical Protein–Protein Interaction. Protein Engineering Design and Selection, 24, 197–207. https://doi.org/10.1093/protein/gzq047
Adzhubei, I.A., Schmidt, S., Peshkin, L., Ramensky, V.E., Gerasimova, A., et al. (2010) A Method and Server for Predicting Damaging Missense Mutations. Nat Methods, 7, 248–249. https://doi.org/10.1038/nmeth0410-248
Li, M.X., Kwan, J.S., Bao, S.Y., Yang, W., Ho, S.L., et al. (2013) Predicting Mendelian Disease-Causing Non-Synonymous Single Nucleotide Variants in Exome Sequencing Studies. PLoS Genet, 9, e1003143. https://doi.org/10.1371/journal.pgen.1003143
Gnad, F., Baucom, A., Mukhyala, K., Manning, G. and Zhang, Z. (2013) As-sessment of Computational Methods for Predicting the Effects of Missense Mutations in Human Cancers. BMC Genomics, 14, S7.
Reva, B., Antipin, Y. and Sander, C. (2011) Predicting the Functional Impact of Protein Mutations: Application to Cancer Genomics. Nucleic Acids Res, 39, e118. https://doi.org/10.1093/nar/gkr407
Dehouck, Y., Kwasigroch, J.M., Rooman, M. and Gilis, D. (2013) BeAtMuSiC: prediction of Changes in Protein-Protein Binding Affinity on Mutations. Nucleic Acids Res, 41, W333–339. https://doi.org/10.1093/nar/gkt450
Kumar, P., Henikoff, S. and Ng, P.C. (2009) Predicting the Effects of Coding Non-Synonymous Variants on Protein Function Using the SIFT Algorithm. Nature Protocols, 4, 1073–1081. https://doi.org/10.1038/nprot.2009.86
Dehouck, Y., Kwasigroch, J.M., Gilis, D. and Rooman, M. (2011) PoPMuSiC 2.1: A Web Server for the Estimation of Protein Stability Changes upon Mutation and Sequence Optimality. BMC Bioinformatics, 12,151. https://doi.org/10.1186/1471-2105-12-151
Adzhubei, I.A., Schmidt, S., Peshkin, L., Ramensky, V.E., Gerasimova, A., Bork, P., et al. (2010) A Method and Server for Predicting Damaging Missense Mutations. Nat Methods, 7, 248–249. https://doi.org/10.1038/nmeth0410-248
Song, Y. X., Zhou, X., Wang, Z., Gao, P., Li, A.L., et al. (2012) The Association be-tween Individual SNPs or Haplotypes of Matrix Metalloproteinase 1 and Gastric Cancer Susceptibility. Progression and Prognosis, 7, e 38002.
Hamosh, A., Scott, A.F. Amberger, J.S., Bocchini, C.A. and McKusick, V.A. (2005) Online Mendelian Inheritance in Man (OMIM), a Knowledgebase of Human Genes and Genetic Disorders. Nucleic Acids Research, 33, D514–D517. https://doi.org/10.1093/nar/gki033
Szklarczyk, D., Franceschini, A., Kuhn, M., Simonovic, M., Roth, A., Minguez, P., et al. (2011) The STRING Database in 2011: Functional Interaction Networks of Proteins, Globally Integrated and Scored. Nucleic Acids Res, 39, D561–8. https://doi.org/10.1093/nar/gkq973
Bairoch, A. and Apweiler, R. (2000) The SWISS-PROT Protein Sequence Database and Its Supplement TrEMBL in 2000. Nucleic Acids Res, 28, 45–48. https://doi.org/10.1093/nar/28.1.45
Schymkowitz, J., Borg, J., Stricher, F., Nys, R., Rousseau, F. and Serrano, L. (2005) The FoldX Web Server: An Online Force Field. Nucl. Acids Res, 33 W382–8. https://doi.org/10.1093/nar/gki387
Shihab, H.A., Gough, J., Cooper, D.N., Day, I.N.M. and Gaunt, T.R. (2013) Predicting the Functional Consequences of Cancer-Associated Amino Acid Substitutions. Bioinformatics, 29,1504-1510. https://doi.org/10.1093/bioinformatics/btt182
Thomas, P.D., Kejariwal, A., Guo, N., Mi, H.Y. and Campbell, M.J., Muruganujan, A. and Lazareva-Ulitsky, B. (2006) Applications for Protein Sequence-Function Evolution Data: mRNA/protein Expression Analysis and Coding SNP Scoring Tools. Nucl. Acids Res, 34, W645-W650. https://doi.org/10.1093/nar/gkl229
Breiman, L. (2001) Random Forests. Machine Learning, 45, 5–32. https://doi.org/10.1023/A:1010933404324
Cortes, C. and Vapnik, V. (1995) Support Vector Machine. Machine Learning, 20, 273-297. https://doi.org/10.1023/A:1022627411411
Stein, A. and Aloy, P. (2010) Novel Peptide-Mediated Interactions Derived from High-Resolution 3-dimensional Structures. PloSComput. Biol, 6, e1000789. https://doi.org/10.1371/journal.pcbi.1000789
Thorisson, G.A., Smith, A.V., Krishnan, L. and Stein, L.D. (2005) The International Hap Map Project Website. Genome Res, 15, 1592-1593. https://doi.org/10.1101/gr.4413105
Kent, W.J., Sugnet, C.W., Furey, T.S., Roskin, K.M., Pringle, T.H., Zahler, A.M. and Haussler, D. (2002) The Human genome browser at UCSC. Genome Res, 12, 996-1006. https://doi.org/10.1101/gr.229102
McVean, et al. (2012) An Integrated Map of Genetic Variation from 1092 Human Genomes. Nature, 491, 56-65. https://doi.org/10.1038/nature11632
Hanley, J. and McNeil, B. (1982) The Meaning and Use of the Area under a Receiver Operating Characteristics (ROC) Curve. Radiology, 143, 29-36. https://doi.org/10.1148/radiology.143.1.7063747