Semi-Global Inference in Phenotype-Protein Network
- 1 Institute of Architecture of Application Systems, University of Stuttgart, Stuttgart, Germany
- 2 School of Computer Science, Harbin Institute of Technology at Weihai, Weihai, China
- 3 Institute of Microelectronics, Chinese Academy of Sciences, Beijing, China
- 4 Department of Computer Science, The University of Hong Kong, Hong Kong, China
Abstract
Discovering genetic basis of diseases is an important goal and a challenging problem in bioinformatics research. In spired by network-based global inference approach, Semi-global inference method is proposed to capture the complex associations between phenotypes and genes. The proposed method integrates phenotype similarities and protein-protein interactions, and it establishes the profile vectors of phenotypes and proteins. Then the relevance between each candi date gene and the target phenotype is evaluated. Candidate genes are then ranked according to relevance mark and genes that are potentially associated with target disease are identified based on this ranking. The model selects nodes in integrated phenotype-protein network for inference, by exploiting Phenotype Similarity Threshold (PST), which throws lights on selection of similar phenotypes for gene prediction problem. Different vector relevance metrics for computing the relevance marks of candidate genes are discussed. The performance of the model is evaluated on Online Mendelian Inheritance in Man (OMIM) data sets and experimental evaluation shows high performance of proposed Semi-global method outperforms existing global inference methods.
- D. Botstein and N. Risch, “Discovering Genotypes Underlying Human Phenotypes: Past Successes for Mendelian Disease, Future Approaches for Complex Disease,” Nature Genetics, Vol. 33, 2003, pp. 228-237. http://dx.doi.org/10.1038/ng1090
- F. S. Turner, D. R. Clutterbuck and C. Semple, “Pocus: Mining Genomic Sequence Annotation to Predict Disease Genes,” Genome Biology, Vol. 4, 2003, p. R75. http://dx.doi.org/10.1186/gb-2003-4-11-r75
- J. Chen, C. Shen and A. Sivachenko, “Mining Alzheimer Disease Relevant Proteins from Integrated Protein Interactome Data,” Pacific Symposium on Biocomputing, Vol. 11, 2006, pp. 367-378.
- A. Hamosh, A. F. Scott, J. S. Amberger, C. A. Bocchini, and V. A. McKusick, “Online Mendelian Inheritance in Man (OMIM), a Knowledge-base of Human Genes and Genetic Disorders,” Nucleic Acids Research, Vol. 33, Database Issue, 2005.
- E. Adie, R. R. Adams, K. L. Evans, D. J. Porteous and B. Pickard, “Speeding Disease Gene Discovery by Sequence Based Candidate Prioritization,” BMC Bioinformatics, Vol. 6, 2005, p. 55. http://dx.doi.org/10.1186/1471-2105-6-55
- L. Sam, Y. Liu, J. Li, C. Friedman and Y. A. Lussier, “Discovery of Protein Interaction Networks Shared by Diseases,” Pacific Symposium on Biocomputing, Vol. 12, 2007, pp. 76-87.
- G. Jiminez-Sanchez, et al., ”Human Disease Genes,” Nature, Vol. 409, 2001, pp. 853-854 http://dx.doi.org/10.1038/35057050
- M. Oti and H. G. Brunner, “The Modular Nature of Genetic Diseases,” Clinical Genetics, Vol. 71, 2007, pp. 1- 11. http://dx.doi.org/10.1111/j.1399-0004.2006.00708.x
- J. H. Jing-Dong, “Understanding Biological Functions through Molecular Networks,” Cell Research, Vol. 18, 2008, pp. 224-237. http://dx.doi.org/10.1111/j.1399-0004.2006.00708.x
- J. Chen, B. Aronow and A. Jegga, “Disease Candidate Gene Identification and Prioritization Using Protein Interaction Networks,” BMC Bioinformatics, Vol. 10, No. 1, 2009, p. 73. http://dx.doi.org/10.1186/1471-2105-10-73
- M. Oti, B. Snel, M. A. Huynen and H. G. Brunner, “Predicting Disease Genes Using Protein-Protein Interactions,” Journal of Medical Genetics, Vol. 43, 2006, pp. 691-698. http://dx.doi.org/10.1136/jmg.2006.041376
- S. Navlakha and C. Kingsford, “The Power of Protein Interaction Networks for Associating Genes with Diseases,” Bioinformatics, Vol. 26, 2010, pp. 1057-1063. http://dx.doi.org/10.1093/bioinformatics/btq076