Improving Disease Prevalence Estimates Using Missing Data Techniques
- 1 Pan African University, Institute for Basic Sciences Technology and Innovation (PAUISTI), Nairobi, Kenya
- 2 Taita Taveta University, Taita Taveta, Kenya
- 3 Universite Gaston Berger de Saint Louis, Saint Louis, Senegal
Abstract
The prevalence of a disease in a population is defined as the proportion of people who are infected. Selection bias in disease prevalence estimates occurs if non-participation in testing is correlated with disease status. Missing data are commonly encountered in most medical research. Unfortunately, they are often neglected or not properly handled during analytic procedures, and this may substantially bias the results of the study, reduce the study power, and lead to invalid conclusions. The goal of this study is to illustrate how to estimate prevalence in the presence of missing data. We consider a case where the variable of interest (response variable) is binary and some of the observations are missing and assume that all the covariates are fully observed. In most cases, the statistic of interest, when faced with binary data is the prevalence. We develop a two stage approach to improve the prevalence estimates; in the first stage, we use the logistic regression model to predict the missing binary observations and then in the second stage we recalculate the prevalence using the observed data and the imputed missing data. Such a model would be of great interest in research studies involving HIV/AIDS in which people usually refuse to donate blood for testing yet they are willing to provide other covariates. The prevalence estimation method is illustrated using simulated data and applied to HIV/AIDS data from the Kenya AIDS Indicator Survey, 2007.
- Hogan, D.R., Salomon, J.A., Canning, D., Hammitt, J.K., Zaslavsky, A.M. and Barnighausen, T. (2012) National HIV Prevalence Estimates for Sub-Saharan Africa: Controlling Selection Bias with Heckman-Type Selection Models. Sexually Transmitted Infections, 88, 17-23.
- Horton, N.J. and Laird, N.M. (2001) Maximum Likelihood Analysis of Logistic Regression Models with Incomplete Covariate Data and Auxiliary Information. Biometrics, 57, 34-42. https://doi.org/10.1111/j.0006-341X.2001.00034.x
- Ibrahim, J.G., Chen, M.-H., Lipsitz, S.R. and Herring, A.H. (2005) Missing-Data Methods for Generalized Linear Models: A Comparative Review. Journal of the American Statistical Association, 100, 332-347. https://doi.org/10.1198/016214504000001844
- Haukoos, J.S. and Newgard, C.D. (2007) Advanced Statistics: Missing Data in Clinical Research-Part1: An Introduction and Conceptual Framework. Academic Emergency Medicine, 14, 662-668.
- Mishra, V., Barrere, B., Hong, R. and Khan, S. (2008) Evaluation of Bias in HIV Seroprevalence Estimates from National Household Surveys. Sexually Transmitted Infections, 84, 63-70.