Different Feature Selection of Soil Attributes Influenced Clustering Performance on Soil Datasets
- 1 School of Urban and Environmental Science, Huaiyin Normal University, Huai’an, China
- 2 Information Engineering School, Nanchang University, Nanchang, China
Abstract
Feature selection is very important to obtain meaningful and interpretive clustering results from a clustering analysis. In the application of soil data clustering, there is a lack of good understanding of the response of clustering performance to different features subsets. In the present paper, we analyzed the performance differences between k -means, fuzzy c -means, and spectral clustering algorithms in the conditions of different feature subsets of soil data sets. The experimental results demonstrated that the performances of spectral clustering algorithm were generally better than those of k -means and fuzzy c -means with different features subsets. The feature subsets containing environmental attributes helped to improve clustering performances better than those having spatial attributes and produced more accurate and meaningful clustering results. Our results demonstrated that combination of spectral clustering algorithm with the feature subsets containing environmental attributes rather than spatial attributes may be a better choice in applications of soil data clustering.
- Jain, A.K., Murty, M.N. and Flynn, P.J. (1999) Data Clustering: A Reviewing. ACM Computing Surveys, 31, 264-323. https://doi.org/10.1145/331499.331504
- Blum, A.L. and Langley, P. (1997) Selection of Relevant Features and Examples in Machine Learning. Artificial Intelligence, 97, 245-271. https://doi.org/10.1016/S0004-3702(97)00063-5
- Guyon, I. and Elisseeff, A. (2003) An Introduction to Variable and Feature Selection. Journal of Machine Learning Research, 3, 1157-1182. https://doi.org/10.1162/153244303322753616
- Xu, R. and Donald, W. (2005) Survey of Clustering Algorithm. IEEE Transactions on Natural Networks, 16, 645-678. https://doi.org/10.1109/TNN.2005.845141
- Young, F.J. and Hammer, R.D. (2000) Defining Geographic Soil Bodies by Landscape Position, Soil Taxonomy and Cluster Analysis. Soil Science Society of America Journal, 64, 948-998. https://doi.org/10.2136/sssaj2000.643989x
- Araujo, S.R., Wetterlind, J., Dematte, J.A.M. and Stenberg, B. (2014) Improving the Prediction Performance of a Large Tropical vis-NIR Spectroscopic Soil Library from Brazil by Clustering into Smaller Subsets or Use of Data Mining Calibration Techniques. European Journal of Soil Science, 65, 718-729. https://doi.org/10.1111/ejss.12165
- Triantafilis, J., Gibbs, I. and Earl, N. (2013) Digital Soil Pattern Recognition in the Lower Namoi Valley Using Numerical Clustering of Gamma-Ray Spectrometry Data. Geoderma, 192, 407-421. https://doi.org/10.1016/j.geoderma.2012.08.021
- Davatgar, N., Neishabouri, M.R. and Sepashhah, A.R. (2012) Delineation of Site Specific Nutrient Management Zones for a Paddy Cultivated Area Based on Soil Fertility Using Fuzzy Clustering. Geoderma, 173-174, 111-118. https://doi.org/10.1016/j.geoderma.2011.12.005
- Tripathi, R., Nayak, A.K., Shahid, M., Lai, B., Gautam, P., Raja, R., Mohanty, S., Kumar, A., Panda, B.B. and Sahoo, P.N. (2015) Delineation of Soil Management Zones for a Rice Cultivated Area in Eastern India Using Fuzzy Clustering. Catena, 133, 128-136. https://doi.org/10.1016/j.catena.2015.05.009
- Odeh, I.O.A., McBratney, A.B. and Chittleborough, D.J. (1990) Design of Optimal Sample Spacing for Mapping Soil Using Fuzzy-k-Means and Regionalized Variable Theory. Geoderma, 47, 93-122. https://doi.org/10.1016/0016-7061(90)90049-F
- Lin, Q.H., Li, H., Luo, W., Lin, Z.M. and Li, B.G. (2013) Optimal Soil Sampling Design for Rubber Tree Management Based on Fuzzy Clustering. Forest Ecology and Management, 308, 214-222. https://doi.org/10.1016/j.foreco.2013.07.028