Recent developments in database technology have seen a wide variety of data being stored in huge collections. The wide variety makes the analysis tasks of a generic database a strenuous task in knowledge discovery. One approach is to summarize large datasets in such a way that the resulting summary dataset is of manageable size. Histogram has received significant attention as summarization/representative object for large database. But, it suffers from computational and space complexity. In this paper, we propose an idea to transform the histogram object into a Piecewise Linear Regression (PLR) line object and suggest that PLR objects can be less computational and storage intensive while compared to th ose of histograms. On the other hand to carry out a cluster analysis, we propose a distance measure for computing the distance between the PLR lines. Case study is presented based on the real data of online education system LMS. This demonstrate s that PLR is a powerful knowledge representative for very large database.
KeywordsHistogramPiecewise Linear RegressionKnowledge DiscoveryBig DataCluster Analysis
Adhikari, A. and Rao, P.R. (2008) Synthesizing Heavy Association Rules from Different Real Data Sources. Pattern Recognition Letters, 29, 59-71. http://dx.doi.org/10.1016/j.patrec.2007.09.001
Fayyad, U., Piatetsky-Shapiro, G. and Smyth, P. (1996) Knowledge Discovery and Data Mining: Towards a Unifying Framework. Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining, Portland, Oregon, 2-4 August 1996, 82-88.
Han, J. and Kamber, M. (2006) Data Mining—Concepts and Techniques. 2nd Edition, Elsevier, Amsterdam.
Piatestsky, S.U., Smyth, M.A. and Uthurusamy, R. (1996) Advances in Knowledge Discovery and Data Mining. AAAI/MIT Press.
Diday, E. (1990) Knowledge Representation and Symbolic Data Analysis. In: Schader, M. and Gaul, W., Eds., Knowledge Data and Computer Assisted Decisions, Springer-Verlag, Berlin, 17-34. http://dx.doi.org/10.1007/978-3-642-84218-4_2
Bock, H. and Diday, E. (2000) Symbolic Objects. In: Bock, H.-H. and Diday, E., Eds., Analysis of Symbolic Data, Springer-Verlag, Berlin, 54-77. http://dx.doi.org/10.1007/978-3-642-57155-8_4
Billard, L. and Diday, E. (2002b) From the Statistics of Data to the Statistics of Knowledge: Symbolic Data Analysis. Journal of the American Statistical Association, 98, 470-487. http://dx.doi.org/10.1198/016214503000242
Lazar, N. (2013) The Big Picture: Symbolic Data Analysis. CHANCE, 26, 39-42. http://dx.doi.org/10.1080/09332480.2013.845450
Billard, L. and Diday, E. (2006) Symbolic Data Analysis. Conceptual Statistics and Data Mining. Wiley, Hoboken. http://dx.doi.org/10.1002/9780470090183
Gioia, F. (2001) Statistical Methods for Interval Variables. Ph.D. Thesis, University Federico II Naples, Naples. (In Italian)
Lauro, C.N. and Palumbo, F. (2000) Principal Component Analysis of Interval Data: A Symbolic Data Analysis Approach. Computational Statistics, 15, 73-87. http://dx.doi.org/10.1007/s001800050038
Gibbs, A.L. and Su, F.E. (2002) On Choosing and Bounding Probability Metrics. International Statistical Review, 70, 419-435. http://dx.doi.org/10.1111/j.1751-5823.2002.tb00178.x
Verde, R. and Irpino, A. (2010) Ordinary Least Squares for Histogram Data Based on Wasserstein Distance. Proceedings of COMPSTAT’2010, Paris, 22-27 August 2010, 581-589. http://dx.doi.org/10.1007/978-3-7908-2604-3_60
Kim, J. and Billard, L. (2013) Dissimilarity Measures for Histogram-Valued Observations, Communications in Statistics-Theory and Method, 42, 283-303. http://dx.doi.org/10.1080/03610926.2011.581785
Pradeep, K.R. and Nagabhushan, P. (2007) An Approach Based on Regression Line Features for Low Complexity Content Based Image Retrieval. International Conference on Computing: Theory and Applications, Kolkata, 5-7 March 2007, 600-604.
Signoriello, S. (2008) Contributions to Symbolic Data Analysis: A Model Data Approach. PhD Thesis, Department of Mathematics and Statistics, University of Naples Federico II, Naples.
Boor, C. (1978) A Practile Guide to Splinea. Springer, New York.
Dias, S. and Brito, P. (2011) A New Linear Regression Model for Histogram-Valued Variables. Proceedings of the 58th ISI World Statistics Congress, Dublin, 21-26 August 2011.
Malash, G.F. and El-Khaiary, M.I. (2010) Piecewise Linear Regression: A Statistical Method for the Analysis of Experimental Adsorption Data by the Intraparticle-Diffusion Models Chemical Engineering Journal, 163, 256-263. http://dx.doi.org/10.1016/j.cej.2010.07.059
Toms, J. and Lesperance, M. (2008) Piecewise Regression: A Tool For Identifying Ecological Thresholds. Ecology, 84, 2034-2041. http://dx.doi.org/10.1890/02-0472
McGee, V.E. and Carleton, W.T. (1970) Piecewise Regression. Journal of the American Statistical Association, 65, 1109-1124. http://dx.doi.org/10.2307/2284278
Wainer, H. (1971) Piecewise Regression: A Simplified Procedure. British Journal of Mathematical and Statistical Psychology, 24, 83-92. http://dx.doi.org/10.1111/j.2044-8317.1971.tb00450.x
Gonzalez, R.C. and Woods, R.E. (2008) Digital Image Processing. 3rd Edition, Prentice Hall, Upper Saddle River.
Bacon, D.W. and Watts, D.G. (1971) Estimating the Transition between Two Intersecting Straight Lines. Biometrika, 58, 525-534. http://dx.doi.org/10.1093/biomet/58.3.525
Shao, J. and Tu, D. (1995) The Jackknife and the Bootstrap. Springer Verlag, New York. http://dx.doi.org/10.1007/978-1-4612-0795-5
Abdi, H. and Williams, L.J. (2010) Jackknife. In: Salkind, N. and Frey, B., Eds., Encyclopedia of Research Design, Sage, Thousand Oaks.
Lavine, B.K. (2012) Clustering and Classification of Analytical Data. In: Encyclopedia of Analytical Chemistry.
Romero, C., Ventura, S. and Garcia, E. (2007) Data Mining in Course Management Systems: Moodle Case Study and Tutorial. Computers and Education, Elsevier Sciences Amsterdam.
http://moodle.org/
Miranda, E. (2011) Data Mining as a Technique to Analyze the Learning Styles of Students in Using the Learning Management System. 2011 Semimar Nasional Aplikasi Teknologi Informasi, 17-18 June 2011.
Jain, A. and Dubes, R. (1988) Algorithms for Clustering Data. Prentice-Hall, Englewood Cliffs.