A novel voting system for the identification of eukaryotic genome promoters
- 1
- 2
- 3
- 4
Abstract
Motivation: Accurate identification and delineation of promoters/TSSs (transcription start sites) is important for improving genome annotation and devising experiments to study and understand transcriptional regulation. Many promoter identifiers are developed for promoter identification. However, each promoter identifier has its own focuses and limitations, and we introduce an integration scheme to combine some identifiers together to gain a better prediction performance. Result: In this contribution, 8 promoter identifiers (Proscan, TSSG, TSSW, FirstEF, eponine, ProSOM, EP3, FPROM) are chosen for the investigation of integration. A feature selection method, called mRMR (Minimum Redundancy Maximum Relevance), is novelly transferred to promoter identifier selection by choosing a group of robust and complementing promoter identifiers. For comparison, four integration methods (SMV, WMV, SMV_IS, WMV_IS), from simple to complex, are developed to process a training dataset with 1400 se- quences and a testing dataset with 378 sequences. As a result, 5 identifiers (FPROM, FirstEF, TSSG, epo- nine, TSSW) are chosen by mRMR, and the integration of them achieves 70.08% and 67.83% correct prediction rates for a training dataset and a testing dataset respectively, which is better than any single identifier in which the best single one only achieves 59.32% and 61.78% for the training dataset and testing dataset respectively.
- Abeel, T., Saeys, Y., Bonnet, E., Rouze, P. and Van de Peer, Y. (2008) Generic eukaryotic core promoter prediction using structural features of DNA. Genome Research, 18(2), 310-323.
- Abeel, T., Saeys, Y., Rouze, P. and Van de Peer, Y. (2008) ProSOM: Core promoter prediction based on unsupervised clustering of DNA physical profiles. Bioinformatics, 24(13), i24-31.
- Davuluri, R.V., Grosse, I. and Zhang, M.Q. (2001) Computational identification of promoters and first exons in the human genome. Nature Genetics, 29(4), 412-417.
- Down, T.A. and Hubbard, T.J. (2002) Computational detection and location of transcription start sites in ma- mmalian genomic DNA. Genome Research, 12(3), 458- 461.
- Prestridge, D.S. (1995) Predicting Pol II promoter sequences using transcription factor binding sites. Journal of Molecular Biology, 249(5), 923-932.
- Solovyev, V.V. and Shahmuradov, I.A. (2003) PromH: Promoters identification using orthologous genomic sequences. Nucleic Acid Research, 31(13), 3540-3545.
- Solovyev, V.V. and Salamov, A. (1997) The Gene-Finder computer tools for analysis of human and model organism genome sequences. The Fifth International Conference on Intelligent Systems for Molecular Biology, 294- 302.
- Werner, T. (1999) Models for prediction and recognition of eukaryotic promoters. Mamm Genome, 10(2), 168- 175.
- Altincay, H. and Demirekler, M. (2000) An information theoretic framework for weight estimation in the com- bination of probabilistic classifiers for speaker identification. Speech Communication, 30(4), 255-272.
- Liu, R. and States, D.J. (2002) Consensus promoter identification in the human genome utilizing expressed gene markers and gene modeling. Genome Research, 12(3), 462-469.
- Lam, L. and Suen, C.Y. (1994) A theoretical-analysis of the application of majority voting to pattern-recognition. 12th IAPR International Conference on Pattern Recognition, Jerusalem, Israel, 418-420.
- Lam, L. and Suen, C.Y. (1997) Application of majority voting to pattern recognition: An analysis of its behavior and performance. IEEE Transactions on Pattern Analysis and Machine Intelligence, 27(5), 553-568.
- Stajniak, A., Szostakowski, J. and Skoneczny, S. (1997) Mixed neural-traditional classifier for character recognition. SPIE-International Society for Optical Engineering, 2949, 102-110.