On Statistical Measures for Data Quality Evaluation
- 1 Department of Public Health Sciences, Henry Ford Health System, Detroit, USA
Abstract
Most GIS databases contain data errors. The quality of the data sources such as traditional paper maps or more recent remote sensing data determines spatial data quality. In the past several decades, different statistical measures have been developed to evaluate data quality for different types of data, such as nominal categorical data, ordinal categorical data and numerical data. Although these methods were originally proposed for medical research or psychological research, they have been widely used to evaluate spatial data quality. In this paper, we first review statistical methods for evaluating data quality, discuss under what conditions we should use them and how to interpret the results, followed by a brief discussion of statistical software and packages that can be used to compute these data quality measures.
- Zeiler, M. (1999) Modeling Our World: The ESRI Guide to Geodatabase Design. ESRI, Inc., New York.
- Devillers, R., Stein, A., Bédard, Y., Chrisman, N., Fisher, P. and Shi, W. (2010) Thirty Years of Research on Spatial Data Quality: Achievements, Failures, and Opportunities. Transactions in GIS, 14, 387-400. https://doi.org/10.1111/j.1467-9671.2010.01212.x
- Fleiss, J.L., Levin, B. and Paik, M.C. (2013) Statistical Methods for Rates and Proportions. John Wiley & Sons, New York.
- Cohen, J. (1960) A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20, 37-46. https://doi.org/10.1177/001316446002000104
- Cohen, J. (1968) Weighted Kappa: Nominal Scale Agreement Provision for Scaled Disagreement or Partial Credit. Psychological Bulletin, 70, 213. https://doi.org/10.1037/h0026256
- Bland, J.M. and Altman, D. (1986) Statistical Methods for Assessing Agreement between Two Methods of Clinical Measurement. The lancet, 327, 307-310. https://doi.org/10.1016/S0140-6736(86)90837-8
- Bartko, J.J. (1966) The Intraclass Correlation Coefficient as a Measure of Reliability. Psychological Reports, 19, 3-11. https://doi.org/10.2466/pr0.1966.19.1.3
- Shrout, P.E. and Fleiss, J.L. (1979) Intraclass Correlations: Uses in Assessing Rater Reliability. Psychological bulletin, 86, 420. https://doi.org/10.1037/0033-2909.86.2.420
- Fielding, A.H. and Bell, J.F. (1997) A Review of Methods for the Assessment of Prediction Errors in Conservation Presence/Absence Models. Environmental Conservation, 24, 38-49. https://doi.org/10.1017/S0376892997000088
- Sinha, S. (2019) Bland-Altman Analysis for Evaluating AHP-Based Wildlife Habitat Suitability Models. Research & Reviews: Journal of Space Science & Technology, 4, 11-18. https://doi.org/10.37591/.v4i2.1958
- Rhew, I.C., Vander Stoep, A., Kearney, A., Smith, N.L. and Dunbar, M.D. (2011) Validation of the Normalized Difference Vegetation Index as a Measure of Neighborhood Greenness. Annals of Epidemiology, 21, 946-952. https://doi.org/10.1016/j.annepidem.2011.09.001
- Feng, X., Wang, Y., Chen, L., Fu, B. and Bai, G. (2010) Modeling Soil Erosion and Its Response to Land-Use Change in Hilly Catchments of the Chinese Loess Plateau. Geomorphology, 118, 239-248. https://doi.org/10.1016/j.geomorph.2010.01.004
- Koo, T.K. and Li, M.Y. (2016) A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine, 15, 155-163. https://doi.org/10.1016/j.jcm.2016.02.012