Confidence Intervals for the Mean of Non-Normal Distribution: Transform or Not to Transform
- 1 Department of Psychology, York University, Toronto, Canada
- 2 Department of Mathematics and Statistics, York University, Toronto, Canada
- 3 Department of Psychology, York University, Toronto, Canada
Abstract
In many areas of applied statistics, confidence intervals for the mean of the population are of interest. Confidence intervals are typically constructed as-suming normality although non-normally distributed data are a common occurrence in practice. Given a large enough sample size, confidence intervals for the mean can be constructed by applying the Central Limit Theorem or by the bootstrap method. Another commonly used method in practice is the back-transformation method, which takes on the following three steps. First, apply a transformation to the data such that the transformed data are normally distributed. Second, obtain confidence intervals for the transformed mean in the usual manner, which assumes normality. Third, apply the back- transformation to obtain confidence intervals for the mean of the original, non-transformed distribution. The parametric Wald method and a small sample likelihood-based third order method, which can address non-normality, are also reviewed in this paper. Our simulation results suggest that common approaches such as back-transformation give erroneous and misleading results even when the sample size is large. However, the likelihood-based third order method gives extremely accurate results even when the sample size is small.
- American Education Research Association (2006) Standards for Reporting on Empirical Social Science Research in AERA Publications. Educational Researcher, 35, 33-40. https://doi.org/10.3102/0013189X035006033
- Cumming, G. (2014) The New Statistics: Why and How. Psychological Science, 25, 7-9. https://doi.org/10.1177/0956797613504966
- Wilkinson, L. and the Task Force on Statistical Inference (1999) Statistical Methods in Psychology Journals: Guidelines and Explanations. American Psychologist, 54, 594-604. https://doi.org/10.1037/0003-066X.54.8.594
- Cumming, G. and Fidler, F. (2009) Confidence Intervals: Better Answers to Better Questions. Journal of Psychology, 217, 15-26. https://doi.org/10.1027/0044-3409.217.1.15
- Cumming, G. and Finch, S. (2001) A Primer on the Understanding, Use, and Calculation of Confidence Intervals that Are Based on Central and Noncentral Distributions. Educational and Psychological Measurement, 61, 532-574. https://doi.org/10.1177/0013164401614002
- Greenland, S., Senn, S.J., Rothman, K.J., Carlin, J.B., Poole, C. Goodman, S.N. and Altman, D.G. (2016) Statistical Tests, P Values, Confidence Intervals, and Power: A Guide to Misinterpretations. European Journal of Epidemiology, 31, 337-350. https://doi.org/10.1007/s10654-016-0149-3
- Moore, D.S., McCabe, G.P. and Craig, B.A. (2014) Introduction to the Practice of Statistics 8th Edition, W.H. Freeman and Company, New York.
- Cain, M.K., Zhang, Z. and Yuan, K.H. (2016) Univariate and Multivariate Skewness and Kurtosis for Measuring Nonnormality: Prevalence, Influence and Estimation. Behavior Research Methods, 1-20. https://doi.org/10.3758/s13428-016-0814-1
- Micceri, T. (1989) The Unicorn, the Normal Curve and Other Improbable Creatures. Psychological Bulletin, 105, 156-166. https://doi.org/10.1037/0033-2909.105.1.156
- Bland, J.M. and. Altman, D.G. (1996) Transformations, means, and confidence intervals. British Medical Journal, 312, 1079. https://doi.org/10.1136/bmj.312.7038.1079
- McDonald, J.H. (2014) Handbook of Biological Statistics. Sparky House, Maryland.
- Efron, B. and Tibshirani, R.J. (1994) An Introduction to the Bootstrap. Chapman and Hall, New York.
- Box, G.E. and Cox, D.R. (1964) An Analysis of Transformation (with Discussion). Journal of the Royal Statistical Society B, 26, 211-252.