Research ArticleOpen AccessGoogle Scholar indexed
Clustering Categorical Data Based on Within-Cluster Relative Mean Difference
School of Mathematics and Statistics, Lanzhou University, Lanzhou, China
School of Mathematics and Statistics, Lanzhou University, Lanzhou, China
- 1 School of Mathematics and Statistics, Lanzhou University, Lanzhou, China
- 2 School of Mathematics and Statistics, Lanzhou University, Lanzhou, China
Open Journal of Statistics·Volume 07 (2017)·Pages 173–181·Published 20 April 2017·DOI10.4236/ojs.2017.72013
Copy link · social · email
Abstract
The clustering on categorical variables has received intensive attention. In dataset with categorical features, some features show the superior performance on clustering procedure. In this paper, we propose a simple method to find such distinctive features by comparing pooled within-cluster mean relative difference and then partition the data upon such features and give subspace of the subgroups. The applications on zoo data and soybean data illustrate the performance of the proposed method.
KeywordsClusteringCategorical VariableDistinctive AttributePooled Within-Cluster Mean Relative DifferenceHamming Distance
- He, Z., Xu, X. and Deng, S. (2008) k-ANMI: A Mutual Information Based Clustering Algorithm for Categorical Data. Information Fusion, 9, 223-233.
- Andritsos, P. and Tsaparas, P. (2010) Categorical Data Clustering. In Sammut, C. and Webb, G., Eds., Encyclopedia of Machine Learning, Springer, Boston, 154-159.
- Bontemps, D. and Toussile, W. (2013) Clustering and Variable Selection for Categorical Multivariate Data. Electronic Journal of Statistics, 7, 2344-2371. https://doi.org/10.1214/13-EJS844
- Anderlucci, L. and Hennig, C. (2014) The Clustering of Categorical Data: A Comparison of a Model-Based and a Distance-Based Approach. Communication in Statistics—Theory and Methods, 43, 704-721. https://doi.org/10.1080/03610926.2013.806665
- Bouguessa, M. (2015) Clustering Categorical Data in Projected Spaces. Data Mining and Knowledge Discovery, 29, 3-38. https://doi.org/10.1007/s10618-013-0336-8
- Silvestre Cardoso, C.M. and Figueiredo, M. (2015) Feature Selection for Clustering Categorical Data with an Embedded Modelling Approach. Expert Systems, 32, 444-453. https://doi.org/10.1111/exsy.12082
- dos Santos, T.R.L. and Zarate, L.E. (2015) Categorical Data Clustering: What Similarity Measure to Recommend? Expert Systems with Applications, 42, 1247-1260.
- Clarke, B.S., Amiri, S. and Clarke, J.L. (2016) EnsCat: Clustering of Categorical Data via Ensembling. BMC Bioinformatics, 17, 380. https://doi.org/10.1186/s12859-016-1245-9
- Ganti, V., Gehrke, J. and Ramakrishnan, R. (1999) CACTUS—Clustering Categorical Data Using Summaries. Proceedings of the 5th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, ACM Press, San Diego, 73-83.
- Guha, S., Rastogi, R. and Shim, K. (2000) Rock: A Robust Clustering Algorithm for Categorical Attributes. Information Systems, 25, 345-366.
- Arias-Castro, E. and Xiao, P. (2017) A Simple Approach to Sparse Clustering. Computational Statistics & Data Analysis, 105, 217-228.
- Zhang, P., Wang, X. and Song, P.X.K. (2006) Clustering Categorical Data Based on Distance Vectors. Journal of the American Statistical Association, 101, 355-367. https://doi.org/10.1198/016214505000000312