A Novel Treatment Optimization System and Top Gene Identification via Machine Learning with Application on Breast Cancer
- 1 Cranbrook Educational Community, Bloomfield Hills, Michigan, USA
- 2 Cranbrook Educational Community, Bloomfield Hills, Michigan, USA
Abstract
Traditional treatment selection of cancers mainly relies on clinical observations and doctor’s judgment, but most outcomes can hardly be predicted. Through Genomics Topology, we use 272 breast cancer patients’ clinical and gene information as an example to propose a treatment optimization and top gene identification system. This study faces certain challenges such as collinearity and the Curse of Dimensionality within data, so by the idea of Analysis of Variance (ANOVA), Principal Component Analysis (PCA) is implemented to resolve this issue. Several genes, for example, SLC40A1 and ACADSB, are found to be both statistically significant and biological-studies supported; the model developed can precisely predict breast cancer mortality, recurrence time, and survival time, with an average MSE of 3.697, accuracy rate of 88.97%, and F1 score of 0.911. The result and methodology used in this study provide a channel for people to further look into the more precise prediction of other cancer outcomes through machine learning and assist in the discovery of targetable pathways for next-generation cancer treatment methods.
- Trop, I., Dugas, A., David, J., El Khoury, M., Boileau, J.F., Larouche, N. and Lalonde, L. (2011) Breast Abscesses: Evidence-Based Algorithms for Diagnosis, Management, and Follow-Up. Radiographics, 31, 1683-1699. https://doi.org/10.1148/rg.316115521
- Edgar, R., Domrachev, M. and Lash, A.E. (2002) Gene Expression Omnibus: NCBI Gene Expression and Hybridization Array Data Repository. Nucleic Acids Research, 30, 207-210. https://doi.org/10.1093/nar/30.1.207
- Ramanan, D. and Angelov, B. (2016) NKI Breast Cancer Data. https://data.world/deviramanan2016/nki-breast-cancerdata
- Kohavi, R. (1995) A Study of cross-validation and Bootstrap for Accuracy Estimation and Model Selection. In Ijcai, 14, 1137-1145.
- Jolliffe, I.T. (1986) Principal Component Analysis and Factor Analysis. In: Principal Component Analysis, Springer, New York, 115-128. https://doi.org/10.1007/978-1-4757-1904-8_7
- Neter, J., Kutner, M.H., Nachtsheim, C.J. and Wasserman, W. (1996) Applied Linear Statistical Models. Vol. 4, Irwin, Chicago, 318.
- Sakamoto, Y., Ishiguro, M. and Kitagawa, G. (1986) Akaike Information Criterion Statistics.
- Akaike, H. (1976) Canonical Correlation Analysis of Time Series and the Use of an Information Criterion. Mathematics in Science and Engineering, 126, 27-96. https://doi.org/10.1016/S0076-5392(08)60869-3
- Tibshirani, R. (1996) Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society, Series B (Methodological), 73, 267-288.
- Cortes, C. and Vapnik, V. (1995) Support Vector Machine. Machine Learning, 20, 273-297. https://doi.org/10.1007/BF00994018
- Breiman, L. (2001) Random Forests. Machine Learning, 45, 5-32. https://doi.org/10.1023/A:1010933404324
- Hosmer Jr., D.W., Lemeshow, S. and Sturdivant, R.X. (2013) Ap-plied Logistic Regression. Vol. 398, John Wiley & Sons, Hoboken.
- McLachlan, G.J. (2004) Discriminant Analysis and Statistical Pattern Recognition. Wiley Interscience, Hoboken.
- Cover, T. and Hart, P. (1967) Nearest Neighbor Pattern Classification. IEEE Transactions on Information Theory, 13, 21-27. https://doi.org/10.1109/TIT.1967.1053964
- Quinlan, J.R. (1987) Simplifying Decision Trees. International Journal of Man-Machine Studies, 27, 221-234. https://doi.org/10.1016/S0020-7373(87)80053-6