Research ArticleOpen AccessGoogle Scholar indexed
Combined Use of k-Mer Numerical Features and Position-Specific Categorical Features in Fixed-Length DNA Sequence Classification
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
- 1 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 2 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 3 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 4 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 5 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 6 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 7 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 8 Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
- 9 Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
Journal of Biomedical Science and Engineering·Volume 10 (2017)·Pages 390–401·Published 26 July 2017·DOI10.4236/jbise.2017.108030
Copy link · social · email
Abstract
To classify DNA sequences, k-mer frequency is widely used since it can convert variable-length sequences into fixed-length and numerical feature vectors. However, in case of fixed-length DNA sequence classification, subsequences starting at a specific position of the given sequence can also be used as categorical features. Through the performance evaluation on six datasets of fixed-length DNA sequences, our algorithm based on the above idea achieved comparable or better performance than other state-of-the art algorithms.
KeywordsSequence ClassificationNumerical and Categorical FeaturesFeature Selection
- GenBank and WGS Statistics. https://www.ncbi.nlm.nih.gov/genbank/statistics/
- UniProt Consortium (2014) UniProt: A Hub for Protein Information. Nucleic Acids Research, 43, D204-D212.
- Xing, Z., Pei, J. and Keogh, E. (2010) A Brief Survey on Sequence Classification. ACM SIGKDD Explorations Newsletter, 12, 40-80. https://doi.org/10.1145/1882471.1882478
- Borozan, I., Watt, S. and Ferretti, V. (2015) Integrating Alignment-Based and Alignment-Free Sequence Similarity Measures for Biological Sequence Classification. Bioinformatics, 31, 1396-1404. https://doi.org/10.1093/bioinformatics/btv006
- Chen, L. and Guo, G. (2014) Nearest Neighbor Classification of Categorical Data by Attributes Weighting. Expert Systems with Applications, 42, 3142-3149. https://doi.org/10.1016/j.eswa.2014.12.002
- Iqbal, M.J., Faye, I., Samir, B.B. and Said, A.M. (2014) Efficient Feature Selection and Classification of Protein Sequence Data in Bioinformatics. The Scientific World Journal, 2014, Article ID: 173869.
- Weitschek, E., Cunial, F. and Felici, G. (2015) LAF: Logic Alignment Free and Its Application to Bacterial Genomes Classification. BioData Mining, 8, 2015. https://doi.org/10.1186/s13040-015-0073-1
- Pham, T.H., Tran, T.B., Ho, T.B., Satou, K. and Valiente, G. (2005) Qualitatively Predicting Acetylation and Methylation Areas in DNA Sequences. Genome Informatics, 16, 3-11.
- Pokholok, D.K., Harbison, C.T., Levine, S., Cole, M., Hannett, N.M., Lee, T.I., Bell, G.W., Walker, K., Rolfe, P.A., Herbolsheimer, E., Zeitlinger, J., Lewitter, F., Gifford, D.K. and Young, R.A. (2005) Genome-Wide Map of Nucleosome Acetylation and Methylation in Yeast. Cell, 122, 517-527. https://doi.org/10.1016/j.cell.2005.06.026
- Higashihara, M., Rebolledo-Mendez, J.D., Yamada, Y. and Satou, K. (2008) Application of a Feature Selection Method to Nucleosome Data: Accuracy Improvement and Comparison with Other Methods. WSEAS Transactions on Biology and Biomedicine, 5, 153-162.
- Li, J. and Wong, L. (2003) Using Rules to Analyse Bio-Medical Data: A Comparison between C4.5 and PCL. Proceedings of Advances in Web-Age Information Management 4th International Conference, Chengdu, 17-19 August, 254-265. https://doi.org/10.1007/978-3-540-45160-0_25
- Nguyen, N.G., Tran, V.A., Ngo, D.L., Phan, D., Lumbanraja, F.R., Faisal, M.R., Abapihi, B., Kubo, M. and Satou, K. (2016) DNA Sequence Classification by Convolutional Neural Network. Journal of Biomedical Science and Engineering, 9, 280-286. https://doi.org/10.4236/jbise.2016.95021