Research ArticleOpen AccessGoogle Scholar indexed
Estimating the number of data clusters via the contrast statistic
Department of Medical Biophysics, Medical Informatics and Biostatistics, National Medical University, Donetsk, Ukraine
Department of Medical Biophysics, Medical Informatics and Biostatistics, National Medical University, Donetsk, Ukraine
Department of Medical Biophysics, Medical Informatics and Biostatistics, National Medical University, Donetsk, Ukraine
Department of Medical Biophysics, Medical Informatics and Biostatistics, National Medical University, Donetsk, Ukraine
- 1 Department of Medical Biophysics, Medical Informatics and Biostatistics, National Medical University, Donetsk, Ukraine
- 2 Department of Medical Biophysics, Medical Informatics and Biostatistics, National Medical University, Donetsk, Ukraine
- 3 Department of Medical Biophysics, Medical Informatics and Biostatistics, National Medical University, Donetsk, Ukraine
- 4 Department of Medical Biophysics, Medical Informatics and Biostatistics, National Medical University, Donetsk, Ukraine
Journal of Biomedical Science and Engineering·Volume 05 (2012)·Pages 95–99·Published 28 February 2012·DOI10.4236/jbise.2012.52012
Copy link · social · email
Abstract
A new method (the Contrast statistic) for estimating the number of clusters in a set of data is proposed. The technique uses the output of self-organising map clustering algorithm, comparing the change in dependency of “Contrast” value upon clusters number to that expected under a uniform distribution. A simulation study shows that the Contrast statistic can be used successfully either, when variables describing the object in a multi-dimensional space are independent (ideal objects) or dependent (real biological objects).
KeywordsSOM Neural NetworkClusteringGap StatisticSilhouette Statistic
- Behbahani, S., Nasrabadi, A. (2009) Application of SOM neural network in clustering. Journal Biomedical Science and Engineering, 2, 637-643. doi:10.4236/jbise.2009.28093
- Kohonen, T. (1982) Self-organized formation of topologically correct feature maps. Biological Cybernetics, 43, 59-69. doi:10.1007/BF00337288
- Tibshirani, R., Walther, G. and Hastie, T. (2000) Estimating the number of cluster in a dataset via the gap statistic. Technical Report, Department of Biostatistics, Stanford University, Stanford.
- Dudoit, S. and Fridlyand, J. (2002) A prediction—based resampling method for estimating the number of clusters in a dataset. Genome Biology, 3, 1-21. doi:10.1186/gb-2002-3-7-research0036
- Sugar, C. and James, G. (2003) Finding the number of clusters in a dataset: An information-theoretic approach. Journal of the American Statistical Association, 98, 750-763. doi:10.1198/016214503000000666
- Tibshirani, R. and Walther, G. (2005) Cluster validation by prediction strength. Journal of Computational & Graphical Statistics, 14, 511-528. doi:10.1198/106186005X59243
- Guo, P., Chen, P. and Lyu, M. (2002) Cluster number selection for a small set of samples using the Bayesian Ying-Yang model. IEEE Transactions on Neural Networks, 13, 757-763. doi:10.1109/TNN.2002.1000144
- Gangnon, R. and Clayton, M. (2007) Cluster detection using Bayes factors from over-parameterized cluster models. Environmental and Ecological Statistics; 14, 69-82. doi:10.1007/s10651-006-0007-7
- Yin, Z., Zhou, X.B., Bakal, C., Li1, F.H., Sun, Y.X., Perrimon, N. and Wong, S.T.C. (2008) Using iterative cluster merging with improved gap statistics to perform online phenotype discovery in the context of high-throughput RNAi screensBMC. Bioinformatics, 9, 264.
- Sharma, A., Podolsky, R., Zhao, J. and McIndoe, R.A. (2009) A modified hyperplane clustering algorithm allows for efficient and accurate clustering of extremely large datasets. Bioinformatics, 25, 1152-1157. doi:10.1093/bioinformatics/btp123
- Medvedovic, M. and Sivaganesan, S. (2002) Bayesian infinite mixture model based clustering of gene expression profiles. Bioinformatics, 18, 1194-1206. doi:10.1093/bioinformatics/18.9.1194
- Qin, Z.S. (2006) Clustering microarray gene expression data using weighted Chinese restaurant process. Bioinformatics, 22, 1988-1997. doi:10.1093/bioinformatics/btl284