An Average Power Primer: Clarifying Misconceptions about Average Power and Replicability

Authors

  • Maria D Soto University of Toronto image/svg+xml
  • Ulrich Schimmack University of Toronto

DOI:

https://doi.org/10.15626/MP.2025.4767

Keywords:

statistical power, replicability, meta-analysis, credibility, uncertainty

Abstract

The replication crisis heightened interest in methods for assessing the credibility of published research. One approach is to estimate the average power of original studies based on observed data. Critics have challenged this approach as an “ontological error”, as a poor predictor of replication outcomes, and as too imprecise to be useful. This article aims to address these critics. We respond by clarifying that using observed data to estimate true power is a standard inferential practice and does not constitute an ontological error. The goal of average power estimation is not to predict the outcome of future replication studies, but the hypothetical outcome if original researchers had to replicate their studies with new samples. Lastly, we demonstrate that even with substantial uncertainty, average power estimates remain diagnostically informative, especially under selection for statistical significance. Using a z-curve analysis of terror management research, we illustrate that a corpus of significant findings does not, by itself, imply evidential value under severe selection for significance. We conclude that, despite limitations, average power estimation remains a valid tool for evaluating published research.

References

Bartoš, F. (2024). The Untrustworthy Evidence in Dishonesty Research. Meta-Psychology, 8. https://doi.org/10.15626/MP.2023.3987

Bartoš, F., & Schimmack, U. (2022). Z-curve 2.0: Estimating Replication Rates and Discovery Rates. Meta-Psychology, 6. https://doi.org/10.15626/MP.2021.2720

Bartoš, F., & Schimmack, U. (2026). Zcurve: An Implementation of Z-Curves (Version 2.4.6). https://cran.r- project.org/web/packages/zcurve/index.html

Brunner, J., & Schimmack, U. (2020). Estimating Population Mean Power Under Conditions of Heterogeneity and Selection for Significance. Meta-Psychology, 4. https://doi.org/10.15626/MP.2018.874

Buckley, J., Hyland, T., & Seery, N. (2023). Estimating the replicability of technology education research. International Journal of Technology and Design Education, 33(4), 1243–1264. https://doi.org/10.1007/s10798-022-09787-6

Chen, L., Benjamin, R., Guo, Y., Lai, A., & Heine, S. J. (2025). Managing the terror of publication bias: A systematic review of the mortality salience hypothesis. Journal of Personality and Social Psychology, 129(1), 20–41. https://doi.org/10.1037/pspa0000438

Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences. L. Erlbaum Associates.

Crede, M., & Sotola, L. K. (2024). All is well that replicates well: The replicability of reported moderation and interaction effects in leading organizational sciences journals. The Journal of Applied Psychology, 109(10), 1659–1667. https://doi.org/10.1037/apl0001197

Doyen, S., Klein, O., Pichon, C.-L., & Cleeremans, A. (2012). Behavioral Priming: It’s All in the Mind, but Whose Mind? PLOS ONE, 7(1), e29081. https://doi.org/10.1371/journal.pone.0029081

Fanelli, D. (2012). Negative results are disappearing from most disciplines and countries. Scientometrics, 90(3), 891–904. https://doi.org/10.1007/s11192-011-0494-7

Francis, G. (2014). The frequency of excess success for articles in Psychological Science. Psychonomic Bulletin & Review, 21(5), 1180–1187. https://doi.org/10.3758/s13423-014-0601-x

Hoenig, J. M., & Heisey, D. M. (2001). The Abuse of Power: The Pervasive Fallacy of Power Calculations for Data Analysis. The American statistician, 55(1), 19–24. https://doi.org/10.1198/000313001300339897

Holgado, D., Mesquida, C., & Román-Caballero, R. (2023). Assessing the Evidential Value of Mental Fatigue and Exercise Research. Sports Medicine, 53(12), 2293–2307. https://doi.org/10.1007/s40279-023-01926-w

Ioannidis, J. P. A., & Trikalinos, T. A. (2007). The appropriateness of asymmetry tests for publication bias in meta-analyses: A large survey. CMAJ: Canadian Medical Association journal = journal de l’Association medicale canadienne, 176(8), 1091–1096. https://doi.org/10.1503/cmaj.060410

John, L., Loewenstein, G., & Prelec, D. (2012). Measuring the Prevalence of Questionable Research Practices With Incentives for Truth Telling. Psychological Science, 23(5), 524–532. https://doi.org/10.1177/0956797611430953

LeBel, E. P., McCarthy, R. J., Earp, B. D., Elson, M., & Vanpaemel, W. (2018). A Unified Framework to Quantify the Credibility of Scientific Findings. Advances in Methods and Practices in Psychological Science, 1(3), 389–402. https://doi.org/10.1177/2515245918787489

Maier, M., Bartoš, F., Oh, M., Wagenmakers, E.-J., Shanks, D., & Harris, A. (2022). Adjusting for Publication Bias Reveals That Evidence for and Size of Construal Level Theory Effects is Substantially Overestimated. https://doi.org/10.31234/osf.io/r8nyu

McAuliffe, W. H. B., Edson, T. C., Louderback, E. R., LaRaja, A., & LaPlante, D. A. (2021). Responsible product design to mitigate excessive gambling: A scoping review and z-curve analysis of replicability. PLOS ONE, 16(4), e0249926. https://doi.org/10.1371/journal.pone.0249926

McShane, B. B., Böckenholt, U., & Hansen, K. T. (2020). Average Power: A Cautionary Note. Advances in Methods and Practices in Psychological Science, 3(2), 185–199. https://doi.org/10.1177/2515245920902370

Mesquida, C., Murphy, J., Lakens, D., & Warne, J. (2023). Publication bias, statistical power and reporting practices in the Journal of Sports Sciences: Potential barriers to replicability. Journal of Sports Sciences, 41(16), 1507–1517. https://doi.org/10.1080/02640414.2023.2269357

Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716

Pek, J., Hoisington-Shaw, K. J., & Wegener, D. T. (2024). Uses of Uncertain Statistical Power: Designing Future Studies, Not Evaluating Completed Studies. Psychological Methods. https://doi.org/10.1037/met0000577

Schimmack, U. (2012). The ironic effect of significant results on the credibility of multiple-study articles. Psychological Methods, 17(4), 551–566. https://doi.org/10.1037/a0029487

Schimmack, U. (2021). The Validation Crisis in Psychology. Meta-Psychology, 5. https://doi.org/10.15626/MP.2019.1645

Schimmack, U. (2025). Z-Curve.3.0 Tutorial: Chapter 2 - Replicability-Index. https://replicationindex.com/2025/07/11/z-curve-3-0-tutorial-chapter-2/

Schimmack, U., & Bartoš, F. (2023). Estimating the false discovery risk of (randomized) clinical trials in medical journals based on published p-values (N. Bobrovitz, Ed.). PLOS ONE, 18(8), e0290084. https://doi.org/10.1371/journal.pone.0290084

Schneck, A. (2023). Are most published research findings false? Trends in statistical power, publication selection bias, and the false discovery rate in psychology (1975–2017) (F. Naudet, Ed.). PLOS ONE, 18(10), e0292717. https://doi.org/10.1371/journal.pone.0292717

Simonsohn, U., Nelson, L. D., & Simmons, J. P. (2014). P-curve: A key to the file-drawer. Journal of Experimental Psychology. General, 143(2), 534–547. https://doi.org/10.1037/a0033242

Sorić, B. (1989). Statistical "Discoveries" and Effect-Size Estimation. Journal of the American Statistical Association, 84(406), 608–610. https://doi.org/10.2307/2289950

Soto, M. D., & Schimmack, U. (2025a). False Positive Psychology? An Empirical Investigation of the False Positive Risk in Psychological Science Before and After the Credibility Revolution. https://doi.org/10.31234/osf.io/6ybeu_v2

Soto, M. D., & Schimmack, U. (2025b). Credibility of results in emotion science: A Z-curve analysis of results in the journals Cognition & Emotion and Emotion. Cognition & Emotion, 39(8), 1803–1819. https://doi.org/10.1080/02699931.2024.2443016

Sotola, L. K. (2023). How Can I Study from Below, that which Is Above? : Comparing Replicability Estimated by Z-Curve to Real Large-Scale Replication Attempts. Meta-Psychology, 7. https://doi.org/10.15626/MP.2022.3299

Sotola, L. K., & Credé, M. (2023). Estimating the replicability of statistically significant moderation effects in personality research using z-curve analysis. Journal of Research in Personality, 107, 104435. https://doi.org/10.1016/j.jrp.2023.104435

Sterling, T. D., Rosenbaum, W. L., & Weinkam, J. J. (1995). Publication Decisions Revisited: The Effect of the Outcome of Statistical Tests on the Decision to Publish and Vice Versa. The American Statistician, 49(1), 108–112. https://doi.org/10.2307/2684823

Sterling, T. D. (1959). Publication Decisions and their Possible Effects on Inferences Drawn from Tests of Significance—or Vice Versa. Journal of the American Statistical Association, 54(285), 30–34. https://doi.org/10.1080/01621459.1959.10501497

van Zwet, E., Gelman, A., Greenland, S., Imbens, G., Schwab, S., & Goodman, S. N. (2023). A New Look at P Values for Randomized Clinical Trials. NEJM Evidence, 3(1), EVIDoa2300003. https://doi.org/10.1056/EVIDoa2300003

Vazire, S. (2018). Implications of the Credibility Revolution for Productivity, Creativity, and Progress. Perspectives on Psychological Science, 13(4), 411–417. https://doi.org/10.1177/1745691617751884

Veen, M. van, Bartoš, F., Sarafoglou, A., Schelvis, R., Bouter, L., & Coenen, P. (2024). Are there indications of publication bias in occupational health research? An examination of the literature and suggestions for future improvements. https://doi.org/10.17605/OSF.IO/WFG7C

Vohs, K. D., Schmeichel, B. J., Lohmann, S., Gronau, Q. F., Finley, A. J., Ainsworth, S. E., Alquist, J. L., Baker, M. D., Brizi, A., Bunyi, A., Butschek, G. J., Campbell, C., Capaldi, J., Cau, C., Chambers, H., Chatzisarantis, N. L. D., Christensen, W. J., Clay, S. L., Curtis, J., . . . Albarracín, D. (2021). A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect. Psychological Science, 32(10), 1566–1581. https://doi.org/10.1177/0956797621989733

Yuan, K.-H., & Maxwell, S. (2005). On the Post Hoc Power in Testing Mean Differences. Journal of Educational and Behavioral Statistics, 30(2), 141–167. https://doi.org/10.3102/10769986030002141

Downloads

Published

2026-06-17

Issue

Section

Commentaries