Research ArticleOpen AccessGoogle Scholar indexed
D-IMPACT: A Data Preprocessing Algorithm to Improve the Performance of Clustering
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
- 1 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 2 Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
- 3 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 4 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 5 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 6 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 7 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 8 Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa, Japan
- 9 Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
- 10 Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
- 11 Institute of Science and Engineering, Kanazawa University, Kanazawa, Japan
Journal of Software Engineering and Applications·Volume 07 (2014)·Pages 639–654·Published 7 July 2014·DOI10.4236/jsea.2014.78059
Copy link · social · email
Abstract
In this study, we propose a data preprocessing algorithm called D-IMPACT inspired by the IMPACT clustering algorithm. D-IMPACT iteratively moves data points based on attraction and density to detect and remove noise and outliers, and separate clusters. Our experimental results on two-dimensional datasets and practical datasets show that this algorithm can produce new datasets such that the performance of the clustering algorithm is improved.
KeywordsAttractionClusteringData PreprocessingDensityShrinking
- Berkhin, P. (2002) Survey of Clustering Data Mining Techniques. Technical Report, Accrue Software, San Jose.
- Murty, M.N., Jain, A.K. and Flynn, P.J. (1999) Data Clustering: A Review. ACM Computing Surveys, 31, 264-323. http://dx.doi.org/10.1145/331499.331504
- Halkidi, M., Batistakis, Y. and Vazirgiannis, M. (2001) On Clustering Validation Techniques. Journal of Intelligent Information Systems, 17, 107-145. http://dx.doi.org/10.1023/A:1012801612483
- Golub, T.R., et al. (1999) Molecular Classification of Cancer: Class Discovery and Class Prediction by Gene Expression Monitoring. Science, 286, 531-537. http://dx.doi.org/10.1126/science.286.5439.531
- Quinn, A. and Tesar, L. (2000) A Survey of Techniques for Preprocessing in High Dimensional Data Clustering. Proceedings of the Cybernetic and Informatics Eurodays.
- Abdi, H. and Williams, L.J. (2010) Principal Component Analysis. Wiley Interdisciplinary Reviews: Computational Statistics, 2, 433-459. http://dx.doi.org/10.1002/wics.101
- Yeung, K.Y. and Ruzzo, W.L. (2001) Principal Component Analysis for Clustering Gene Expression Data. Bioinformatics, 17, 763-774. http://dx.doi.org/10.1093/bioinformatics/17.9.763
- Shi, Y., Song, Y. and Zhang, A. (2005) A Shrinking-Based Clustering Approach for Multidimensional Data. IEEE Transaction on Knowledge Data Engineering, 17, 1389-1403. http://dx.doi.org/10.1109/TKDE.2005.157
- Chang, F., Qiu, W. and Zamar, R.H. (2007) CLUES: A Non-Parametric Clustering Method Based on Local Shrinking. Computational Statistics & Data Analysis, 52, 286-298. http://dx.doi.org/10.1016/j.csda.2006.12.016
- Jain, A.K. and Dubes, R.C. (1988) Algorithms for Clustering Data. Prentice Hall, Upper Saddle River.
- Ester, M., Kriegel, H.P., Sander, J. and Xu, X. (1996) A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise. Proceedings of the 2nd International Conference on Knowledge Discovery and Data mining, 226-231.
- Ankerst, M., Breunig, M.M., Kriegel, H.P. and Sander, J. (1999) OPTICS: Ordering Points to Identify Clustering Structure. Proceedings of the ACM SIGMOD Conference, 49-60.
- Hinneburg, A. and Keim, D. (1998) An Efficient Approach to Clustering in Large Multimedia Databases with Noise. Proceeding 4th International Conference on Knowledge Discovery & Data Mining, 58-65.