The Sense and Nonsense of Effect Size: What Does Funder and Ozer’s (2019) Call Mean for Health Research?

The Position and Central Thesis of the Article

David C. Funder and Daniel J. Ozer’s article “Evaluating Effect Size in Psychological Research: Sense and Nonsense,” published in Advances in Methods and Practices in Psychological Science, exposes two long-standing misconceptions in the evaluation of effect size in psychological research and proposes two concrete interpretive axes in their place: benchmarking and consequence-oriented reasoning. After demonstrating the arbitrariness of Cohen’s (1988) small–medium–large labels and the misleading character of the “variance explained” rhetoric of r², the authors close with a bold recommendation: in psychology, even an r of .05 should be considered “small in the short term but potentially consequential in the long run,” whereas effect sizes of r ≥ .40 are inflated values that rarely replicate in large samples. This is an epistemic warning that directly concerns health management, clinical epidemiology, and health policy research as well.

The Misleading Nature of r² and the Arbitrariness of Cohen’s Thresholds

Funder and Ozer’s most critical intervention is dismantling the habit of squaring the correlation coefficient to produce a “percentage of variance explained.” Drawing on Darlington’s (1990) coin-tossing example, the authors show how r² is misleading: the correlation between flipping a nickel and the payoff is r = .4472, while the correlation for the dime is r = .8944. When squared, the nickel appears to explain 20% of the variance and the dime 80%, suggesting that the dime matters four times as much as the nickel. The factual reality, however, is that the dime is only twice as valuable as the nickel; the r scale captures this correctly while the r² scale distorts it. Mischel’s (1968) rhetorical framing that “a personality coefficient of .30 explains only 9% of the variance” has, for half a century, fed exactly the kind of computational trick that has made personality psychology look weak. Cohen’s thresholds are, by his own later admission, provisional labels offered “only when no better basis was available”; in the actual distribution derived by Gignac and Szodorai (2016) from 708 meta-analytic correlations, the mean is r = .19 and the upper quartile is r = .29. In other words, the real world of personality and social psychology never even reaches what Cohen called large.

The Logic of Cumulation in Small Effects

The authors’ adoption of Abelson’s (1985) “baseball batting average” analogy forms the epistemic backbone of the article. The correlation between batting skill and the outcome of a single at bat is r = .056; this number seems negligible at first glance. Yet because a player has approximately 550 at bats per season, the difference between a .200 and a .300 average translates into millions of dollars in salary. The authors extend this logic to health outcomes in a straightforward way: the correlation between aspirin use after a heart attack and the prevention of a subsequent attack is on the order of r = .03, yet in a sample of 10,845 individuals this corresponds to 85 lives saved. This implies that, in the evaluation of health policy and public health interventions, neither r nor p alone is sufficient; the cumulation of effects must also be taken into account.

The New Thresholds

Assuming the effect size is reliably estimated, the authors propose the following recalibrated framework: r = .05 is very small at the level of single events but potentially consequential in the long run; r = .10 is still small but with substantial cumulative potential; r = .20 is a medium-sized effect with practical relevance in both the short and long run; r = .30 is a powerful effect; and r ≥ .40 should be treated with skepticism in psychology, and most likely in health behavior research as well, as an inflated and difficult-to-replicate effect. Unlike the conventions of clinical research, this scale is anchored to the actual distribution of effects observed in behavioral science; however, its adoption has profound implications for health management and health economics research as well.

Implications for Health Research and a Critical Appraisal

The framework offered by Funder and Ozer provides a substantial conceptual ground for interpreting effect sizes in the field of health management; at the same time, the article carries serious limitations from the perspective of health research that must be made explicit. These limitations are not merely a historical record but issues that, as of 2026, remain open and will continue to shape the research agenda of the coming period.

The Gap Between Statistical and Clinical Significance

The first significant gap in the article is that it conducts the entire “significance” debate within a statistical and behavioral axis while ignoring the literatures of clinical significance and Minimal Clinically Important Difference (MCID). Even if a health intervention shows an effect of r = .20, whether it crosses the meaningful threshold on the patient’s quality of life, symptom score, or functional measure is a separate question. From the standpoint of the 2026 research agenda, integrating the calibration of behavioral effect sizes with MCID, Patient-Reported Outcome Measures (PROM), and the Patient Acceptable Symptom State remains an open methodological gap. In health management, the effect size debate cannot remain purely statistical or purely clinical; new metrics that bridge the two are required.

Heterogeneous Treatment Effects and the Average Effect Fallacy

The authors acknowledge — though only in a footnote referring to Gelman (2018) — that an average r = .08 effect of a growth mindset intervention may correspond to a one-grade-point change in 10% of students and no change at all in the rest. However, the article leaves this point at the level of a footnote and never engages with the literature on Heterogeneous Treatment Effects (HTE). In health research, the average effect size has increasingly become a contested summary statistic, and Individual Treatment Effects are now estimated through machine-learning–based methods such as causal forests, BART, and meta-learners. In the 2026 research agenda, exposing the unequal distributions hidden beneath the rhetoric of “small effects” and standardizing distributional heterogeneity in effect size reporting are critical priorities.

The Inadequacy of Standardized Effect Sizes for Health Economics and Decision Analysis

The authors offer a sound critique of how standardized effect sizes (r, Cohen’s d) confound consistency with magnitude, arguing that raw regression coefficients are more informative. Yet this critique is not extended to health economics metrics. In health research, decision-relevant metrics — relative risk (RR), odds ratio (OR), hazard ratio (HR), and especially Number Needed to Treat (NNT) and Number Needed to Harm — represent both absolute and interpretable forms of effect size. Funder and Ozer cite the Harding Center for Risk Literacy’s “fact boxes” approvingly, but only in half a paragraph at the end of the article. From the perspective of the 2026 agenda, absolute effect metrics in the logic of NNT should be made a standard reporting component for behavioral and psychological interventions as well; this is becoming an increasingly normative requirement in health policy decision-making.

The Bayesian Framework and Effect Size Uncertainty

The article is written entirely within a frequentist framework. It warns that the 95% confidence interval for r is wide in small samples, but never mentions updating effect size estimates with prior distributions, posterior estimates, Bayes factors, or the Region of Practical Equivalence (ROPE). Yet, since the work of Lakens (2017) and Kruschke (2018), Bayesian interpretations of effect size have rapidly diffused into behavioral and clinical research. As of 2026, how hierarchical Bayesian models that combine small effects with prior knowledge can be integrated into the standard for effect size evaluation in health management research remains an open question, and the Bayesian counterparts of the new thresholds proposed by Funder and Ozer have yet to be established.

The Effect Size Problem in the Era of Big Data and Artificial Intelligence

Funder and Ozer’s recommendation for large samples represented a reform call as of 2019; however, in the era of electronic health records, administrative health databases, and real-world evidence in 2026, the signal of the problem has reversed. Sample sizes now easily exceed n = 100,000 or n = 1,000,000, and p values produce significance with little resistance. Under such conditions, filtering out effect sizes that are statistically significant but lack clinical or managerial meaning must be addressed through the Smallest Effect Size of Interest (SESOI) framework and equivalence testing. The article cites Lakens, Scheel, and Isager (2018) only in a single sentence; from the standpoint of the 2026 research agenda, reporting each health management study with a pre-specified SESOI is now becoming a methodological minimum standard.

Causal Inference and the Structural Interpretation of Effect Size

The most fundamental weakness of the article is that it treats effect size as an associational concept and does not connect it to the modern framework of causal inference. The interpretation of effect sizes computed within Pearl’s do-calculus framework, directed acyclic graphs (DAGs), propensity score matching, instrumental variables, and regression discontinuity is conceptually distinct from Funder and Ozer’s framework. In health policy and health services research, an intervention’s r = .15 effect is interpreted within a causal frame as the Average Treatment Effect; this should not be conflated with the observational r found in personality psychology. In the 2026 agenda, the calibration of effect sizes must be reconstructed in a manner sensitive to the type of causal estimand involved (ATT, ATE, CATE, LATE).

The Absence of Post-Replication-Crisis Tools

The article references the replication crisis and advocates for large-sample preregistered studies; however, it does not engage with post-2019 meta-analytic bias-correction techniques such as p-curve, z-curve, R-Index, TIVA, and PET-PEESE, nor with publication-bias–corrected effect sizes in particular. In health management, the majority of systematic reviews still assess publication bias only through funnel plots and Egger’s test. From the standpoint of the 2026 agenda, bias-resistant effect size estimators and the tracking of the share of preregistered studies (the registered reports percentage) in health management research should be adopted as indicators of the field’s methodological maturity.

Effect Size in Multilevel and Longitudinal Designs

The article largely operates within the logic of bivariate correlation. Discussion of effect size for multilevel models, repeated-measures designs, and longitudinal structural equation models is virtually absent; yet health management research is woven through cluster data at the organizational level, repeated measurement at the patient level, and hospital–clinician–patient hierarchies. In such designs, Cohen’s f², pseudo-R², and ICC-based effect metrics must be interpreted differently. From the 2026 agenda perspective, developing standard effect size reporting protocols for multilevel designs remains an open task.

The Absence of an Equity Framework

Finally, while Funder and Ozer’s framework positions effect size as a tool of average reasoning, it does not account for inequality and equity. In health inequality research, an intervention’s having different effect sizes across groups (a gradient) is itself the research question; inequality metrics — concentration index, Wagstaff decomposition, Theil index — carry both the directional and distributional components of effect size simultaneously. As a 2026 research agenda item, how the framework of behavioral effect size will speak to health inequality metrics is a priority open area in the health policy research of many countries, including Türkiye.

Conclusion: Toward Responsible Interpretation in Knowledge Production

Funder and Ozer’s (2019) call is a corrective framework for health management and behavioral health research as well: squaring correlations to belittle them, performing ritual small–medium–large reasoning with Cohen’s thresholds, and unquestioningly believing inflated effect numbers must cease to be academic habits. The new thresholds proposed by the authors open the way for the effect sizes routinely encountered in research on health behaviors, health communication, healthcare worker well-being, and service quality (the .10–.25 range) to be interpreted as respectable and meaningful through cumulation. However, because the article does not adequately connect to the frameworks of clinical significance, heterogeneous treatment effects, health economics metrics, Bayesian interpretation, big data, and causal inference, it remains, as of 2026, a starting point rather than a destination. The task of the health management researcher is to preserve Funder and Ozer’s call for honesty while integrating effect size with the components of patient-level meaningfulness, distributional equity, and causal interpretation. Without this integration, effect size reform stands as an achievement in psychology; in health, it remains a half-finished epistemic transformation.

Reference: Funder, D. C., & Ozer, D. J. (2019). Evaluating effect size in psychological research: Sense and nonsense. Advances in Methods and Practices in Psychological Science, 2(2), 156–168. doi:10.1177/2515245919847202 (with the 2020 corrigendum).

Subscribe to the Health Topics Newsletter!

Google reCaptcha: Invalid site key.