Effect Size
A standardized measure of how large a difference or relationship is, independent of sample size — the quantity that tells you whether a statistically significant result is actually meaningful.
What it means
Effect size is a family of statistics that quantify the magnitude of a difference, association, or relationship in standardized units — independent of sample size. Where a p-value tells you whether an effect is probably real, an effect size tells you how big it is.
Common effect size measures:
| Measure | Use case | Small | Medium | Large |
|---|---|---|---|---|
| Cohen’s d | Comparing two means | 0.2 | 0.5 | 0.8 |
| Pearson r | Correlation | 0.1 | 0.3 | 0.5 |
| η² (eta-squared) | ANOVA | 0.01 | 0.06 | 0.14 |
| Odds ratio | Binary outcomes | 1.5 | 2.5 | 4.0+ |
| Relative risk | Incidence rates | 1.2 | 1.5 | 2.0+ |
Cohen’s benchmarks (small/medium/large) are rough guidelines for the behavioral sciences; what constitutes a “large” effect varies enormously across fields — a small effect on a common disease outcome affecting millions of people may be more important than a large effect in a narrow laboratory setting.
Why p-values alone are not enough
With a large enough sample, any effect — however tiny — will produce p < 0.05. An RCT with 100,000 participants might find that an intervention improves a symptom score by 0.3 points on a 100-point scale (p < 0.001, Cohen’s d = 0.03). The result is highly statistically significant; the effect is clinically irrelevant.
Effect size solves this: it gives the result’s practical importance directly, allowing readers to judge whether the finding is worth acting on.
Effect size in meta-analysis
Meta-analyses pool effect sizes across studies to estimate a common underlying effect and its uncertainty. The forest plot — the standard visualization of a meta-analysis — shows the effect size and confidence interval for each study alongside the pooled estimate.
AI-assisted meta-analysis tools like Elicit help extract effect sizes from primary papers. The extracted values should always be verified against the source: effect size reporting conventions vary across fields (some report Cohen’s d, others partial η², others simply means and SDs), and AI extraction can confuse these.