Comparison of Outlier Techniques Based on Simulated Data
- 1 Department of Statistics, Faculty of Physical Sciences, Nnamdi Azikiwe University, Awka, Nigeria
- 2 Monetary & Policy Department, Central Bank of Nigeria, Abuja, Nigeria
- 3 Department of Statistics, Faculty of Physical Sciences, Nnamdi Azikiwe University, Awka, Nigeria 2Monetary & Policy Department, Central Bank of Nigeria, Abuja, Nigeria
Abstract
This research work employed a simulation study to evaluate six outlier techniques: t -Statistic, Modified Z -Statistic, Cancer Outlier Profile Analysis (COPA), Outlier Sum-Statistic (OS), Outlier Robust t -Statistic (ORT), and the Truncated Outlier Robust t -Statistic (TORT) with the aim of determining the technique that has a higher power of detecting and handling outliers in terms of their P -values, true positives, false positives, False Discovery Rate (FDR) and their corresponding Receiver Operating Characteristic (ROC) curves. From the result of the analysis, it was revealed that OS was the best technique followed by COPA, t , ORT, TORT and Z respectively in terms of their P -values. The result of the False Discovery Rate (FDR) shows that OS is the best technique followed by COPA, t , ORT, TORT and Z . In terms of their ROC curves, t -Statistic and OS have the largest Area under the ROC Curve (AUC) which indicates better sensitivity and specificity and is more significant followed by COPA and ORT with the equal significant AUC while Z and TORT have the least AUC which is not significant.
- [1] Grubbs, F.E. (1969) Procedures for Detecting Outlying Observations in Samples. Technometrics, 11, 1-21. http://dx.doi.org/10.1080/00401706.1969.10490657
- Hawkins, D. (1980) Identification of Outliers. Chapman and Hall, Kluwer Academic Publishers, Boston/Dordrecht/ London.
- Aggarwal, C.C. (2005) On Abnormality Detection in Spuriously Populated Data Streams. SIAM Conference on Data Mining. Kluwer Academic Publishers Boston/Dordrech/London.
- Barnett, V. and Lewis, T. (1994) Outliers in Statistical Data. 3rd Edition, John Wiley & Sons, Kluwer Academic Publishers, Boston/Dordrecht/London.
- Dudoit, S., Yang, Y., Callow, M. and Speed, T. (2002) Statistical Methods for Identifying Differentially Expressed Genes in Replicated DNA Microarray Experiments. Statistica Sinica, 12, 111-139.
- Troyanskaya, O.G., Garber, M.E., Brown, P.O., Botstein, D. and Altman, R.B. (2002) Nonparametric Methods for Identifying Differentially Expressed Genes in Microarray Data. Bioinformatics, 18, 1454-1461. http://dx.doi.org/10.1093/bioinformatics/18.11.1454
- Tomlins, S., Rhodes, D., Perner, S., Dhanasekaran, S., Mehra, R., Sun, X., Varambally, S., Cao, X., Tchinda, J., Kuefer, R., et al. (2005) Recurrent Fusion of TMPRSS2 and ETS Transcription Factor Genes in Prostate Cancer. Science, 310, 644-648. http://dx.doi.org/10.1126/science.1117679
- Efron, B., Tibshirani, R., Storey, J. and Tusher, V. (2001) Empirical Bayes Analysis of a Microarray Experiment. Journal of the American Statistical Association, 96, 1151-1160. http://dx.doi.org/10.1198/016214501753382129
- Iglewicz, B. and Hoaglin, D.C. (2010) Detection of Outliers. Engineering Statistical Handbook, 1.3.5.17. Database Systems Group.
- Lyons-Weiler, J., Patel, S., Becich, M. and Godfrey, T. (2004) Tests for Finding Complex Patterns of Differential Expression in Cancers: Towards Individualized Medicine. Bioinformatics, 5, 1-9.
- Tibshirani, R. and Hastie, R. (2006) Outlier Sums Statistic for Differential Gene Expression Analysis. Biostatistics, 8, 2-8. http://dx.doi.org/10.1093/biostatistics/kxl005
- Benjamini, Y. and Hochberg, Y. (1995) Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society Series B, 57, 289-300.
- Wu, B. (2007) Cancer Outlier Differential Gene Expression Detection. Biostatistics, 8, 566-575. http://dx.doi.org/10.1093/biostatistics/kxl029