Systematic Reviews and Meta-Analysis
Systematic reviews use explicit, reproducible methods to identify, appraise, and synthesise all relevant research on a specific question, with meta-analysis providing a quantitative pooled estimate of effect that sits at the apex of the evidence hierarchy.
Key Facts
- Systematic reviews are the highest level of evidence in the evidence hierarchy
- They use explicit, reproducible methodology to identify, select, appraise, and synthesise all relevant studies
- Meta-analysis is the statistical pooling of results from multiple studies to generate a single summary effect estimate
- Forest plot: graphical display showing individual study effects and pooled estimate; diamond represents overall effect
- Heterogeneity assessed by I² statistic: 0-25% low, 25-50% moderate, 50-75% substantial, >75% considerable
- Funnel plot: assesses publication bias; asymmetry suggests bias (small negative studies less likely published)
- PRISMA statement is the reporting standard for systematic reviews and meta-analyses
- Cochrane Library is the largest repository of systematic reviews of healthcare interventions
Overview
Key Facts
Systematic reviews and meta-analyses provide the most reliable evidence for clinical decision-making when conducted rigorously. They synthesise evidence from multiple studies, increasing statistical power and precision.
Steps in Conducting a Systematic Review
- Define research question (PICO format)
- Develop protocol (register on PROSPERO)
- Systematic literature search (multiple databases: MEDLINE, Embase, CENTRAL)
- Study selection (screening by ≥2 reviewers, inclusion/exclusion criteria)
- Data extraction (standardised forms)
- Quality assessment (risk of bias tool)
- Data synthesis (narrative and/or meta-analysis)
- Assessment of certainty of evidence (GRADE)
- Reporting (PRISMA)
Meta-Analysis Methods
- Fixed-effect model: assumes all studies estimate the same underlying effect; appropriate when heterogeneity is low
- Random-effects model: assumes true effects vary between studies; accounts for between-study heterogeneity; more conservative
- Mantel-Haenszel method: common fixed-effect method for dichotomous outcomes
- DerSimonian and Laird: common random-effects method
Assessing Quality
- Risk of bias: assessed for each included study (Cochrane RoB 2 for RCTs, ROBINS-I for non-randomised)
- Heterogeneity: clinical (population/intervention differences), methodological (design differences), statistical (I² and Chi² test)
- Publication bias: funnel plot asymmetry, Egger test, trim and fill method
Clinical Presentation
Interpreting Forest Plots
- Each study represented by a square (size proportional to weight) and horizontal line (95% CI)
- Diamond at bottom represents pooled estimate (width = CI)
- Vertical line at RR/OR = 1 (line of no effect)
- If diamond does not cross line of no effect: statistically significant
- I² value and p-value for heterogeneity reported
Interpreting Funnel Plots
- Plots study effect size (x-axis) against study precision (y-axis, usually SE)
- Symmetric funnel shape expected in absence of publication bias
- Asymmetry (missing small negative studies) suggests publication bias
Subgroup and Sensitivity Analyses
- Subgroup analysis: explores whether effect varies by patient/study characteristics
- Sensitivity analysis: tests robustness (e.g. excluding high risk-of-bias studies)
- Meta-regression: explores sources of heterogeneity using study-level covariates
Differential Diagnosis
| Feature | Systematic Review | Narrative Review | Meta-Analysis |
|---|---|---|---|
| Methods | Explicit, reproducible | Author-selected, subjective | Statistical pooling |
| Search | Systematic, comprehensive | Non-systematic | Part of systematic review |
| Quality assessment | Formal (risk of bias tools) | Informal/none | Included studies assessed |
| Synthesis | Narrative ± quantitative | Narrative only | Quantitative pooled estimate |
| Bias risk | Minimised | High (selection bias) | Depends on included studies |
| Reproducibility | High | Low | High |
Diagnosis / Investigation
Critical Appraisal (CASP Systematic Review Checklist)
- Did the review address a clearly focused question?
- Did the authors look for the right type of papers?
- Were all important relevant studies included?
- Did the authors assess the quality of included studies?
- If results were combined, was it reasonable to do so?
- What is the overall result?
- How precise are the results?
- Can results be applied to the local population?
- Were all important outcomes considered?
- Are the benefits worth the harms and costs?
Key Statistical Concepts
- I²: percentage of variability due to heterogeneity rather than chance; >50% = substantial
- Chi² (Q) test: tests whether observed differences in results are compatible with chance alone
- Prediction interval: range of effects expected in a new study (wider than CI of pooled estimate)
- GRADE certainty: High, Moderate, Low, Very Low
Management
Using Systematic Reviews in Practice
- Check Cochrane Library for existing reviews on clinical questions
- Assess GRADE certainty of evidence before applying to patients
- Consider applicability to your patient population
- NNT from meta-analysis may be the most reliable estimate available
- Living systematic reviews: continuously updated as new evidence emerges
Limitations
- Only as good as included studies ('garbage in, garbage out')
- Publication bias: negative studies less likely published
- Heterogeneity may limit pooling
- Ecological fallacy: pooled results may not apply to individual patients
- Time lag: may not include most recent evidence
Prognosis
- Systematic reviews and meta-analyses form the basis of clinical guidelines (NICE, WHO, Cochrane)
- They provide the most precise estimates of treatment effects through pooled analysis
- Updated/living reviews improve timeliness of evidence synthesis
- Network meta-analyses (comparing multiple treatments simultaneously) increasingly used for guideline development
- Individual patient data meta-analyses provide the most detailed evidence but require data sharing
Other Relevant Information
I² Heterogeneity Interpretation
| I² Value | Heterogeneity Level | Action |
|---|---|---|
| 0-25% | Low | Fixed-effect model appropriate |
| 25-50% | Moderate | Consider random-effects |
| 50-75% | Substantial | Random-effects; explore sources |
| >75% | Considerable | May not be appropriate to pool; explore/explain |
GRADE Certainty of Evidence
| Level | Meaning |
|---|---|
| High | Very confident effect estimate is close to true effect |
| Moderate | Moderately confident; true effect likely close to estimate |
| Low | Limited confidence; true effect may be substantially different |
| Very Low | Very little confidence; true effect likely substantially different |