Hybrid clustering combines partitional and hierarchical clustering for computational effectiveness and versatility in cluster shape. In such clustering, a dissimilarity measure plays a crucial role in the hierarchical merging. The dissimilarity measure has great impact on the final clustering, and data-independent properties are needed to choose the right dissimilarity measure for the problem at hand. Properties for distance-based dissimilarity measures have been studied for decades, but properties for density-based dissimilarity measures have so far received little attention. Here, we propose six data-independent properties to evaluate density-based dissimilarity measures associated with hybrid clustering, regarding equality, orthogonality, symmetry, outlier and noise observations, and light-tailed models for heavy-tailed clusters. The significance of the properties is investigated, and we study some well-known dissimilarity measures based on Shannon entropy, misclassification rate, Bhattacharyya distance and Kullback-Leibler divergence with respect to the proposed properties. As none of them satisfy all the proposed properties, we introduce a new dissimilarity measure based on the Kullback-Leibler information and show that it satisfies all proposed properties. The effect of the proposed properties is also illustrated on several real and simulated data sets.
Hastie, T., Tibshirani, R. and Friedman, J. (2009) The Elements of Statistical Learning: Data mining, Inference, and Prediction. 2nd Edition, Springer Series in Statistics, Springer, New York.
Jain, A.K., Murty, M.N. and Flynn, P.J. (1999) Data Clustering: A Review. ACM Computing Surveys, 31, 264-323. http://dx.doi.org/10.1145/331499.331504
Jain, A.K. (2010) Data Clustering: 50 Years beyond K-Means. Pattern Recognition Letters, 31, 651-666. http://dx.doi.org/10.1016/j.patrec.2009.09.011
Cox, D.R. (1957) Note on Grouping. Journal of the American Statistical Association, 52, 543-547. http://dx.doi.org/10.1080/01621459.1957.10501411
Fisher, W.D. (1958) On Grouping for Maximum Homogeneity. Journal of the American Statistical Association, 53, 789-798. http://dx.doi.org/10.1080/01621459.1958.10501479
Wong, M.A. (1982) A Hybrid Clustering Method for Identifying High-Density Clusters. Journal of the American Statistical Association, 77, 841-847. http://dx.doi.org/10.1080/01621459.1982.10477896
Goldberger, J. and Roweis, S.T. (2005) Hierarchical Clustering of a Mixture Model. Advances in Neural Information Processing Systems.
Lin, C.-R. and Chen, M.-S. (2005) Combining Partitional and Hierarchical Algorithms for Robust and Efficient Data Clustering with Cohesion Self-Merging. IEEE Transactions on Knowledge and Data Engineering, 17, 145-159. http://dx.doi.org/10.1109/TKDE.2005.21
Liu, M., Jiang, X. and Kot, A.C. (2009) A Multi-Prototype Clustering Algorithm. Pattern Recognition, 42, 689-698. http://dx.doi.org/10.1016/j.patcog.2008.09.015
Murty, M.N. and Krishna, G. (1981) A Hybrid Clustering Procedure for Concentric and Chain-Like Clusters. International Journal of Computer & Information Sciences, 10, 397-412. http://dx.doi.org/10.1007/BF00996137
Patra, B.K., Nandi, S. and Viswanath, P. (2011) A Distance Based Clustering Method for Arbitrary Shaped Clusters in Large Datasets. Pattern Recognition, 44, 2862-2870. http://dx.doi.org/10.1016/j.patcog.2011.04.027
Vijaya, P.A., Murty, N.M. and Subramanian, D.K. (2006) Efficient Bottom-Up Hybrid Hierarchical Clustering Techniques for Protein Sequence Classification. Pattern Recognition, 39, 2344-2355. http://dx.doi.org/10.1016/j.patcog.2005.12.001
Viswanath, P. and Suresh Babu, V. (2009) Rough-DBSCAN: A Fast Hybrid Density Based Clustering Method for Large Data Sets. Pattern Recognition Letters, 30, 1477-1488. http://dx.doi.org/10.1016/j.patrec.2009.08.008
Zhang, T., Ramakrishnan, R. and Livny, M. (1996) BIRCH: An Efficient Data Clustering Method for Very Large Databases. Proceedings of the 1996 ACM SIGMOD International Conference on Management of Data, New York, 103-114. http://dx.doi.org/10.1145/233269.233324
Zhong, S. and Ghosh, J. (2003) A Unified Framework for Model-Based Clustering. Journal of Machine Learning Research, 4, 1001-1037.
García-Escudero, L.A., Gordaliza, A., Matrán, C. and Mayo-Iscar, A. (2011) Exploring the Number of Groups in Robust Model-Based Clustering. Statistics and Computing, 21, 585-599. http://dx.doi.org/10.1007/s11222-010-9194-z
Arbelaitz, O., Gurrutxaga, I., Muguerza, J., Pérez, J.M. and Perona, I. (2013) An Extensive Comparative Study of Cluster Validity Indices. Pattern Recognition, 46, 243-256. http://dx.doi.org/10.1016/j.patcog.2012.07.021
Dubes, R. and Jain, A.K. (1976) Clustering Techniques: The User’s Dilemma. Pattern Recognition, 8, 247-260. http://dx.doi.org/10.1016/0031-3203(76)90045-5
Kleinberg, J. (2003) An Impossibility Theorem for Clustering. Advances in Neural Information Processing Systems (NIPS 2002), 15, 463-470.
Puzicha, J., Hofmann, T. and Buhmann, J.M. (2000) A Theory of Proximity Based Clustering: Structure Detection by Optimization. Pattern Recognition, 33, 617-634. http://dx.doi.org/10.1016/S0031-3203(99)00076-X
Puzicha, J., Buhmann, J.M., Rubner, Y. and Tomasi, C. (1999) Empirical Evaluation of Dissimilarity Measures for Color and Texture. The Proceedings of the Seventh IEEE International Conference on Computer Vision, 2, 1165-1172. http://dx.doi.org/10.1109/ICCV.1999.790412
Baudry, J.-P., Raftery, A.E., Celeux, G., Lo, K. and Gottardo, R. (2010) Combining Mixture Components for Clustering. Journal of Computational and Graphical Statistics, 9, 332-353. http://dx.doi.org/10.1198/jcgs.2010.08111
Karypis, G., Han, E.-H. and Kumar, V. (1999) Chameleon: Hierarchical Clustering Using Dynamic Modeling. Computer, 32, 68-75. http://dx.doi.org/10.1109/2.781637
Huber, M.F., Bailey, T., Durrant-Whyte, H. and Hanebeck, U.D. (2008) On Entropy Approximation for Gaussian Mixture Random Vectors. IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems, August 2008, 181-188. http://dx.doi.org/10.1109/MFI.2008.4648062
Maz’ya, V. and Schmidt, G. (1996) On Approximate Approximations Using Gaussian Kernels. IMA Journal of Numerical Analysis, 13-29. http://dx.doi.org/10.1093/imanum/16.1.13
Dempster, A.P., Laird, N.M. and Rubin, D.B. (1977) Maximum Likelihood from Incomplete Data via the EM Algorithm. Journal of the Royal Statistical Society, Series B (Methodological), 39, 1-38.
Schwarz, G. (1978) Estimating the Dimension of a Model. The Annals of Statistics, 6, 461-464. http://dx.doi.org/10.1214/aos/1176344136
Akaike, H. (1974) A New Look at the Statistical Model Identification. IEEE Transactions on Automatic Control, 19, 716-723. http://dx.doi.org/10.1109/TAC.1974.1100705
McLachlan, G. and Peel, D. (2000) Finite Mixture Models. Wiley Series in Probability and Statistics, John Wiley & Sons, Inc. http://dx.doi.org/10.1002/0471721182
Bache, K. and Lichman, M. (2013) UCI Machine Learning Repository.
Franczak, B.C., Browne, R.P. and McNicholas, P.D. (2014) Mixtures of Shifted Asymmetric Laplace Distributions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36, 1149-1157. http://dx.doi.org/10.1109/TPAMI.2013.216
Peel, D. and McLachlan, G.J. (2000) Robust Mixture Modelling Using the t Distribution. Statistics and Computing, 10, 339-348. http://dx.doi.org/10.1023/A:1008981510081
R. Tibshirani, G. Walther, and T. Hastie, \Estimating the number of clusters in a data set via the gap statistic," Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol. 63, pp. 411{423, Jan. 2001.
Biernacki, C., Celeux, G. and Govaert, G. (2000) Assessing a Mixture Model for Clustering with the Integrated Completed Likelihood. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22, 719-725. http://dx.doi.org/10.1109/34.865189
Ali, S.M. and Silvey, S.D. (1966) A General Class of Coefficients of Divergence of One Distribution from Another. Journal of the Royal Statistical Society, Series B (Methodological), 28, 131-142.
Calderero, F. and Marques, F. (2008) General Region Merging Approaches Based on Information Theory Statistical Measures. 15th IEEE International Conference on Image Processing, 3016-3019.
Basseville, M. (1989) Distance Measures for Signal Processing and Pattern Recognition. Signal Processing, 18, 349-369. http://dx.doi.org/10.1016/0165-1684(89)90079-0
Csiszár, I. (1967) Information-Type Measures of Difference of Probability Distributions and Indirect Observations. Studia Scientiarum Mathematicarum Hungarica, 2, 299-318.
Pardo, L. (2006) Statistical Inference Based on Divergence Measures. Vol. 185 of Statistics, Chapman and Hall/CRC.
Berger, A. (1953) On Orthogonal Probability Measures. Proceedings of the American Mathematical Society, 4, 800-806. http://dx.doi.org/10.1090/S0002-9939-1953-0056868-5
Farcomeni, A. (2013) Robust Constrained Clustering in Presence of Entry-Wise Outliers. Technometrics, 56, 102-111. http://dx.doi.org/10.1080/00401706.2013.826148
Huber, P.J. and Ronchetti, E.M. (2009) Robust Statistics. Wiley Series in Probability and Statistics. 2nd Edition, John Wiley & Sons, Inc., Hoboken.
Shannon, C.E. (1948) A Mathematical Theory of Communication, Part I. The Bell System Technical Journal, 27, 379-423. http://dx.doi.org/10.1002/j.1538-7305.1948.tb01338.x
Lin, J. (1991) Divergence Measures Based on the Shannon Entropy. IEEE Transactions on Information Theory, 37, 145-151. http://dx.doi.org/10.1109/18.61115
Bhattacharyya, A. (1943) On a Measure of Divergence between Two Statistical Populations Defined by Their Probability Distributions. Bulletin of the Calcutta Mathematical Society, 35, 99-109.
Kailath, T. (1967) The Divergence and Bhattacharyya Distance Measures in Signal Selection. IEEE Transactions on Communication Technology, 15, 52-60. http://dx.doi.org/10.1109/TCOM.1967.1089532
Kullback, S. and Leibler, R.A. (1951) On Information and Sufficiency. The Annals of Mathematical Statistics, 22, 79-86. http://dx.doi.org/10.1214/aoms/1177729694
Fraley, C. and Raftery, A.E. (2002) Model-Based Clustering, Discriminant Analysis, and Density Estimation. Journal of the American Statistical Association, 97, 611-631. http://dx.doi.org/10.1198/016214502760047131
Mangasarian, O.L., Street, W.N. and Wolberg, W.H. (1995) Breast Cancer Diagnosis and Prognosis via Linear Programming. Operations Research, 43, 570-577. http://dx.doi.org/10.1287/opre.43.4.570
Furman, E. (2008) On a Multivariate Gamma Distribution. Statistics & Probability Letters, 78, 2353-2360. http://dx.doi.org/10.1016/j.spl.2008.02.012