An important problem with null hypothesis significance testing, as it is normally performed, is that it is uninformative to reject a point null hypothesis [1]. A way around this problem is to use range null hypotheses [2]. But the use of range null hypotheses also is problematic. Aside from the usual issues of whether null hypothesis significance tests can be justified at all, there is an issue that is specific to range null hypotheses. It is not straightforward how to calculate the probability of the data given a range null hypothesis. The traditional way is to use the single point that maximizes the obtained p-value. The Bayesian alternative is to propose a prior probability distribution and integrate across it. Because frequentists and Bayesians disagree about a variety of issues, especially those pertaining to whether it is permissible to assign probabilities to hypotheses, and what gets lost in the shuffle is that the two camps actually come to different answers for the probability of the data given a range null hypothesis. Because the probability of the data given the hypothesis is a precursor for both camps, for drawing conclusions about hypotheses, different values for this probability for the different camps is crucial but seldom acknowledged. The goal of the present article is to bring out the problem in a manner accessible to researchers without strong mathematical or statistical backgrounds.
Meehl, P.E. (1967) Theory-Testing in Psychology and Physics: A Methodological Paradox. Philosophy of Science, 34, 103-115. http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.693.8918&rep=rep1&type=pdf https://doi.org/10.1086/288135
Leventhal, L. (1999) Answering Two Criticisms of Hypothesis Testing. Psychological Reports, 85, 3-18. https://doi.org/10.2466/pr0.1999.85.1.3
Neyman, J. and Pearson, E.S. (1928) On the Use and Interpretation of Certain Test Criteria for Purposes of Statistical Inference. Biometrika, 20A, 175-240.
Neyman, J. and Pearson, E.S. (1933) The Testing of Statistical Hypotheses in Relation to Probabilities a Priori. Proceedings of the Cambridge Philosophical Society, 29, 492-510. https://doi.org/10.1017/S030500410001152X
Etz, A. and Vandekerckhove, J. (2016) A Bayesian Perspective on the Reproducibility Project: Psychology. PLoS ONE, 11, e0149794.
Bakan, D. (1966) The Test of Significance in Psychological Research. Psychological Bulletin, 66, 423-437. http://www.tc.umn.edu/~nydic001/docs/teaching/Fall2011_PSY3801H/readings/Readings%20-%2003Bakan%201966.pdf https://doi.org/10.1037/h0020412
Carver, R.P. (1978) The Case against Statistical Significance Testing. Harvard Educational Review, 48, 378-399. http://healthyinfluence.com/wordpress/wp-content/uploads/2015/04/Carver-SSD-1978.pdf https://doi.org/10.17763/haer.48.3.t490261645281841
Carver, R.P. (1993) The Case against Statistical Significance Testing, Revisited. The Journal of Experimental Education, 61, 287-292. http://www.jstor.org/stable/20152382 https://doi.org/10.1080/00220973.1993.10806591
Cohen, J. (1994) The Earth Is Round (p http://qpsy.snu.ac.kr/teaching/gradstat/Cohen.pdf https://doi.org/10.1037/0003-066X.49.12.997
Fisher, R.A. (1973) Statistical Methods and Scientific Inference. 3rd Edition, Hafner Press., New York.
Kass, R.E. and Raftery, A.E. (1995) Bayes Factors. Journal of the American Statistical Association, 90, 773-795. https://www.stat.washington.edu/raftery/Research/PDF/kass1995.pdf https://doi.org/10.1080/01621459.1995.10476572
Meehl, P.E. (1978) Theoretical Risks and Tabular Asterisks: Sir Karl, Sir Ronald, and the Slow Progress of Soft Psychology. Journal of Consulting and Clinical Psychology, 46, 806-834. http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.200.7648&rep=rep1&type=pdf https://doi.org/10.1037/0022-006X.46.4.806
Meehl, P.E. (1990) Appraising and Amending Theories: The Strategy of Lakatosian Defense and Two Principles That Warrant Using It. Psychological Inquiry, 1, 108-141. https://pdfs.semanticscholar.org/2a38/1d2b9ae7e7905a907ad42ab3b7e2d3480423.pdf https://doi.org/10.1207/s15327965pli0102_1
Meehl, P.E. (1997) The Problem Is Epistemology, Not Statistics: Replace Significance Tests by Confidence Intervals and Quantify Accuracy of Risky Numerical Predictions. In: Harlow, L., Mulaik, S.A. and Steiger, J.H., Eds., What If There Were No Significance Tests? Erlbaum, Mahwah, NJ, 393-425.
Rozeboom, W.W. (1960) The Fallacy of the Null-Hypothesis Significance Test. Psychological Bulletin, 57, 416-428. http://www.ufrgs.br/psico-laboratorio/textos_classicos_9.pdf https://doi.org/10.1037/h0042040
Rozeboom, W.W. (1997) Good Science Is Abductive, Not Hypothetico-Deductive. In: Harlow, L., Mulaik, S.A. and Steiger, J.H., Eds., What If There Were No Significance Tests? Erlbaum, Mahwah, NJ, 335-391.
Schmidt, F.L. (1996) Statistical Significance Testing and Cumulative Knowledge in Psychology: Implications for the Training of Researchers. Psychological Methods, 1, 115-129. http://qpsy.snu.ac.kr/teaching/stat_dir/Schmidt(1996).pdf https://doi.org/10.1037/1082-989X.1.2.115
Schmidt, F.L. and Hunter, J.E. (1997) Eight Objections to the Discontinuation of Significance Testing in the Analysis of Research Data. In: Harlow, L., Mulaik, S.A. and Steiger, J.H., Eds., What If There Were No Significance Tests? Erlbaum, Mahwah, NJ, 37-64.
Trafimow, D. (2003) Hypothesis Testing and Theory Evaluation at the Boundaries: Surprising Insights from Bayes’s Theorem. Psychological Review, 110, 526-535. https://doi.org/10.1037/0033-295X.110.3.526
Trafimow, D. (2006) Using Epistemic Ratios to Evaluate Hypotheses: An Imprecision Penalty for Imprecise Hypotheses. Genetic, Social, and General Psychology Monographs, 132, 431-462. https://doi.org/10.3200/MONO.132.4.431-462
Trafimow, D. and Marks, M. (2015) Editorial. Basic and Applied Social Psychology, 37, 1-2. https://doi.org/10.1080/01973533.2015.1012991
Trafimow, D. and Marks, M. (2016) Editorial. Basic and Applied Social Psychology, 38, 1-2. https://doi.org/10.1080/01973533.2016.1141030
Valentine, J.C., Aloe, A.M. and Lau, T.S. (2015) Life after NHST: How to Describe Your Data without “p-ing” Everywhere. Basic and Applied Social Psychology, 37, 260-273. https://doi.org/10.1080/01973533.2015.1060240
Hagen, R.L. (1997) In Praise of the Null Hypothesis Significance Test. American Psychologist, 52, 15-24. https://doi.org/10.1037/0003-066X.52.1.15
Greenwald, A.G. (1975) Consequences of Prejudice against the Null Hypothesis. Psychological Bulletin, 82, 1-20. https://doi.org/10.1037/h0076157
Serlin, R.C. and Lapsley, D.K. (1985) Rationality in Psychological Research. American Psychologist, 40, 73-83. https://pdfs.semanticscholar.org/0ecb/48d1ad3747b4dd78ddcf2c8fd1f546481aa4.pdf https://doi.org/10.1037/0003-066X.40.1.73
Serlin, R.C. and Lapsley, D.K. (1993) Rational Appraisal of Psychological Research and the Good-Enough Principle. In: Keren, G. and Lewis, C., Eds., A Handbook for Data Analysis in the Behavioral Sciences: Methodological Issues, Erlbaum, Hillsdale, NJ, 199-228.
Hays, W.L. (1994) Statistics. 5th Edition, Harcourt Brace College Publishers, Fort Worth, TX.
Howson, C. (1990) Fitting Your Theory to the Facts: Probably Not Such a Bad Thing after All. In: Savage, C.W., Ed., Minnesota Studies in the Philosophy of Science, Vol. 14, University of Minnesota Press, Minneapolis, 224-244.
Howson, C. and Urbach, P. (1989) Scientific Reasoning: The Bayesian Approach. Open Court, La Salle, IL.
Howson, C. and Urbach, P. (1994) Probability, Uncertainty and the Practice of Statistics. In: Wright, G. and Ayton, P., Eds., Subjective Probability, Wiley, Chichester, England, 39-51.
Kruschke, J.K. (2011) Doing Bayesian Data Analysis: A Tutorial with R and BUGS. Academic Press, USA.
Fisher, R.A. (1925) Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh, Scotland.
Fisher, R.A. (1930) Inverse Probability. Proceedings of the Cambridge Philosophical Society, 26, 528-538. https://doi.org/10.1017/S0305004100016297
Wasserstein, R.L. and Lazar, N.A. (2016) The ASA’s Statement on p-Values: Context, Process, and Purpose. The American Statistician, 70, 129-133. https://doi.org/10.1080/00031305.2016.1154108
Vieland, V.J. and Hodge, S.E. (2011) Measurement of Evidence and Evidence of Measurement. Statistical Applications in Genetics and Molecular Biology, 10, Article 35. https://doi.org/10.2202/1544-6115.1682
Trafimow, D. and Earp, B.D. (2017) Null Hypothesis Significance Testing and the Use of P Values to Control the Type I Error Rate: The Domain Problem. New Ideas in Psychology, 45, 19-27. https://doi.org/10.1016/j.newideapsych.2017.01.002
Lakatos, I. (1978) The Methodology of Scientific Research Programmes. Cambridge University Press, Cambridge, UK. https://doi.org/10.1017/CBO9780511621123