Systematic reviews with AIPRA

A chapter-by-chapter guide to conducting a systematic review and how AIPRA supports each stage.

Chapter 7 of 8

Meta-analysis: effect sizes, heterogeneity, models, and forest plots

4 min read · Updated

In short

Key takeaways

  • Studies must be converted to a common effect metric before pooling; use SMD or Hedges' g for continuous outcomes on differing scales, and OR or RR for binary outcomes.
  • Always extract sample size and a measure of precision such as the standard error — meta-analysis weights studies through those variances, so pooling is not defensible without them.
  • I-squared describes the fraction of total variation attributable to heterogeneity rather than within-study error. It is not a measure of how far effects differ on the outcome scale.
  • As a rule of thumb, I-squared of roughly 0–40% is often treated as low and above about 75% as considerable, but these are guidance rather than hard rules.
  • Random-effects models are the default for most reviews; fixed-effect models are appropriate only when studies are nearly identical in design, population, and conduct.
  • If heterogeneity is high, explore it with prespecified subgroup analyses or meta-regression rather than reporting a single pooled effect alone.
  • Funnel plot asymmetry can suggest small-study effects or publication bias, but many non-bias mechanisms distort funnels; interpret cautiously, especially with few studies.

This chapter assumes you have already decided that a quantitative pooled summary is appropriate for at least one outcome—see the Evidence synthesis chapter for when meta-analysis is justified. What follows is a concise tour of the mechanics: effect sizes, heterogeneity, model choice, and how you present and scrutinize results.

Phase 1 — Building blocks: effect sizes

Before you can pool results, studies need to be on a common scale. If one study reports a mean difference and another reports Cohen’s d, you face a “Tower of Babel” problem: the numbers are not yet commensurable (Guilera et al., 2022). Harmonization usually means converting or back-calculating to a standard effect metric, ideally with documented formulas and sensitivity checks.

Continuous outcomes

Use a standardized mean difference (SMD) or Hedges’ g when instruments or scales differ across trials (for example different depression questionnaires). That puts mean changes on a comparable footing when raw units are not the same.

Binary outcomes

Use odds ratios (OR) or relative risks (RR), consistent with your question, study designs, and reporting standards. Be explicit about which contrast you modeled (e.g. event vs. non-event definitions).

Phase 2 — Signal vs. noise: heterogeneity

Meta-analysis is not only about the average effect; it is about whether averaging is coherent. Heterogeneity describes how much effect sizes scatter beyond what sampling error alone would predict. A large pooled estimate means little if studies are effectively answering different questions.

Cochran’s Q

Tests whether there is more between-study dispersion than expected by chance under a fixed-effect framing. It is useful but insensitive when there are few studies or low power.

statistic

Describes approximately what fraction of total variation in point estimates is due to heterogeneity rather than within-study error (not a measure of how far effects differ on the outcome scale).

= ((Q − df) / Q) × 100%

Interpretation (rule-of-thumb): Values of from about 0% to 40% are often treated as low heterogeneity, while values above roughly 75% are often described as considerable (Higgins et al., 2003). Treat these thresholds as guidance, not hard rules: clinical and methodological differences between studies matter as much as the number.

Phase 3 — Choosing your model

The pooling model encodes what you believe about how true effects vary across studies—this is one of the most consequential analysis choices in the review.

ModelAssumptionTypical use
Fixed-effectOne shared “true” effect underlies all studies; observed differences reflect sampling error only.Uncommon in contemporary reviews. Consider only when studies are nearly identical in design, population, and conduct—rare outside specialized settings.
Random-effectsTrue effects differ across studies (e.g. settings, populations, interventions). The model estimates both within-study and between-study variance (DerSimonian & Laird, 1986; many implementations now use refined estimators of τ²).Default for most reviews—it reflects that each study estimates its own context-specific effect while still allowing an average (and prediction intervals where appropriate).

Phase 4 — Forest plot and bias detection

The forest plot is the central figure for meta-analysis: each study’s effect and confidence interval is drawn on a common axis, with a diamond (or similar) summarizing the pooled estimate when you combine studies.

Visual inference: For ratios (OR, RR), a null effect is often drawn at 1.0; for mean differences and SMD, at 0. If the pooled diamond does not cross that null line, the summary is conventionally statistically significant at the chosen alpha—always pair that visual with the numeric estimate, CI width, and heterogeneity statistics.

Publication bias and small-study effects: A funnel plot plots effect size against precision; asymmetry can suggest missing small or “negative” studies, though many non-bias mechanisms also distort funnels. Complement plots with formal tests where appropriate (for example Egger’s regression test for funnel asymmetry in some settings) and interpret results cautiously—especially with few studies.

Frequently asked questions

What does I-squared mean in meta-analysis?

I-squared estimates the percentage of total variation across studies that is due to genuine heterogeneity rather than within-study sampling error. It is calculated as ((Q − df) / Q) × 100%. It does not indicate how large the differences in effect are on the outcome scale.

What is a high I-squared value?

Commonly cited rules of thumb treat roughly 0–40% as low heterogeneity and above about 75% as considerable. These thresholds are guidance only; clinical and methodological differences between studies matter as much as the number, and heterogeneity should be interpreted in context.

Should I use a fixed-effect or random-effects model?

Random-effects is the default for most systematic reviews because it assumes true effects vary across studies and estimates both within-study and between-study variance. Fixed-effect assumes one shared true effect with differences due only to sampling error, which is rarely plausible outside highly standardized settings.

How do you read a forest plot?

Each study appears as its effect estimate with a confidence interval on a common axis, and a diamond summarizes the pooled estimate. The null line is at 1.0 for ratio measures such as odds ratios and relative risks, and at 0 for mean differences and standardized mean differences. If the pooled diamond does not cross the null line, the summary is conventionally statistically significant, but the visual should always be paired with the numeric estimate, interval width, and heterogeneity statistics.

What is a funnel plot used for?

A funnel plot graphs effect size against precision to look for asymmetry that may indicate missing small or negative studies. Asymmetry is suggestive of publication bias or small-study effects but has other possible causes, so it is often complemented by formal tests such as Egger's regression and interpreted with caution when few studies are available.

Which effect size should I use for continuous outcomes?

Use a standardized mean difference or Hedges' g when studies measure the same construct with different instruments or scales, such as different depression questionnaires. A raw mean difference is appropriate only when all studies report the same units.

References

  1. Higgins JPT, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta-analyses. BMJ. 2003;327(7414):557-560. doi:10.1136/bmj.327.7414.557
  2. DerSimonian R, Laird N. Meta-analysis in clinical trials. Controlled Clinical Trials. 1986;7(3):177-188. doi:10.1016/0197-2456(86)90046-2

How to cite this chapter

AIPRA. Meta-analysis: effect sizes, heterogeneity, models, and forest plots. In: The Systematic Review E-book. Updated 2026-09-05.

A plain-text version of this chapter is available at /systematic-review-ebook/meta-analysis.md.