Sensitivity Analysis in Systematic Reviews: Methods, Examples, and Reporting
Sensitivity analysis tests how robust your systematic review findings are to methodological decisions. Learn leave-one-out, threshold, and decision-node approaches with reporting examples.
Dr. Sarah Mitchell
March 4, 2026
Key Takeaways
Sensitivity analysis tests whether the conclusions of a systematic review change when key methodological decisions are varied, distinguishing robust findings from fragile ones
Leave-one-out analysis removes each study in turn to identify influential studies that disproportionately drive the pooled result
Decision-node sensitivity analysis systematically varies choices made at each stage: study inclusion criteria, effect size selection, model specification, and risk of bias thresholds
Sensitivity analyses should be pre-specified in the protocol rather than conducted post hoc, though exploratory sensitivity analyses are acceptable if clearly labeled
PRISMA 2020 requires reporting of all pre-specified sensitivity analyses and their results, regardless of whether findings changed
A finding that is sensitive to reasonable methodological choices is not necessarily wrong, but the uncertainty should be communicated transparently to decision-makers
Sensitivity analysis in systematic reviews tests whether the conclusions of your review hold up when key methodological decisions are varied. Every systematic review involves judgment calls, from which studies to include to which statistical model to use, and sensitivity analysis reveals which of those decisions actually matter for the final result. A finding that survives multiple sensitivity analyses is robust. One that flips under reasonable alternative choices is fragile, and readers deserve to know.
The Cochrane Handbook describes sensitivity analysis as a "crucial component" of systematic reviews, and deep dive into prisma 2020 requires that all pre-specified sensitivity analyses and their results be reported regardless of outcome. Yet many published reviews either skip sensitivity analysis entirely or bury a single leave-one-out analysis in supplementary materials. This guide covers the full toolkit: when sensitivity analysis is needed, which methods to use, how to interpret and report results, and how to pre-specify analyses in your our guide to developing a systematic review protocol.
The core question of sensitivity analysis is simple: "Would my conclusion change if I had made a different reasonable decision?" This applies to every stage of a systematic review:
Study inclusion. Would results differ if borderline studies (unclear eligibility, conference abstracts, unpublished data) were included or excluded?
Data extraction. When a study reports multiple time points, outcome measures, or subgroups, does the choice of which data to extract affect the pooled result?
Risk of bias. Does restricting the analysis to studies with low risk of bias change the conclusion?
Missing data. When studies have incomplete outcome data, do best-case and worst-case imputation scenarios produce different conclusions?
Effect size measure. For binary outcomes, do odds ratios, risk ratios, and risk differences tell the same story?
Each of these represents a decision node where an alternative choice was equally defensible.
Leave-One-Out Analysis
Leave-one-out sensitivity analysis is the most common and most straightforward method. It sequentially removes each study from the meta-analysis for researchers, recalculates the pooled estimate, and examines whether any single study disproportionately influences the result.
How to interpret: If the pooled effect size and its statistical significance remain stable regardless of which study is removed, your findings are robust to individual study influence. If removing a single study changes the direction of the effect (e.g., from favoring treatment to favoring control) or changes statistical significance (from significant to non-significant or vice versa), that study is influential and warrants close examination.
What to do with influential studies: An influential study is not necessarily problematic. It may be the largest, highest-quality study that legitimately carries more weight. Investigate whether it differs clinically (different population, dose, or comparator), methodologically (different design, lower risk of bias), or statistically (different follow-up duration, different outcome definition). Report your findings transparently rather than excluding the study without justification.
Limitations: Leave-one-out analysis only tests single-study influence. It does not detect situations where two or three studies collectively drive the result, nor does it address methodological decisions beyond study inclusion.
Software implementation: In R, metafor::leave1out() performs this automatically. In Stata, metainf provides similar functionality. RevMan does not include built-in leave-one-out analysis. Our online sensitivity analysis tool provides an interactive interface for exploring study influence.
Decision-Node Sensitivity Analysis
Decision-node analysis systematically varies choices made at each stage of the review. Unlike leave-one-out (which only tests study inclusion), this approach examines the full range of methodological decisions.
Pre-specify decision nodes in your protocol. For each node, identify the primary analysis choice and at least one reasonable alternative:
Decision Node
Primary Analysis
Sensitivity Analysis
Study eligibility
Include randomized controlled trials and quasi-experimental
Restrict to randomized controlled trials only
Risk of bias
Include all studies
Restrict to low risk of bias only
Missing data
Complete case analysis
Best-case/worst-case imputation
Statistical model
Random-effects (REML)
Fixed-effect model
Effect measure
Standardized mean difference
Mean difference (if scales comparable)
Outlier handling
Include all studies
Exclude statistical outliers (> 3 SD from pooled mean)
Publication type
Include only peer-reviewed
Add grey literature
Run the meta-analysis under each alternative specification and present results side by side. This gives readers and guideline panels a comprehensive picture of evidence robustness.
Need help with your meta-analysis?
Our PhD statisticians run complete meta-analyses: effect sizes, forest plots, heterogeneity testing, and publication-ready results sections.
Threshold analysis asks: "How much would the data need to change to overturn the conclusion?" Rather than testing specific alternative decisions, it quantifies the fragility of the result.
Fragility index for meta-analysis: For binary outcomes, the fragility index counts the minimum number of events that, if reassigned from treatment to control (or vice versa) across studies, would change the statistical significance of the pooled result. A fragility index of 2 means that reassigning just 2 events would flip the conclusion, indicating a fragile finding.
Threshold for clinical relevance: Beyond statistical significance, you can calculate how much the pooled effect would need to shift to cross a clinically meaningful threshold. If the pooled risk ratio is 0.72 and the minimally important difference is 0.85, the question becomes: "What would need to change for the effect to become clinically unimportant?"
Unmeasured confounding sensitivity analysis: For systematic reviews of observational studies, the E-value quantifies how strong an unmeasured confounder would need to be to explain away the observed association. A large E-value means the result is robust to potential confounding; a small E-value means even weak confounding could account for the finding.
Risk of Bias Sensitivity Analysis
Risk-of-bias sensitivity analysis: walkthrough
Restricting the meta-analysis to studies assessed as having low risk of bias is one of the most important and commonly performed sensitivity analyses. The Cochrane Handbook and GRADE certainty assessment approach both recommend this approach.
Interpreting discordance: If the pooled effect is significant when all studies are included but non-significant when restricted to low-bias studies, this has direct implications for GRADE certainty ratings. The evidence may be rated down for risk of bias if the result depends on studies with serious methodological limitations.
Stratified analysis: Rather than a binary include/exclude approach, stratify studies by risk of bias level (low, some concerns, high) and test for interaction. This reveals whether effect sizes differ systematically by study quality, a pattern sometimes called small-study effects when combined with explore publication bias.
Reporting Sensitivity Analysis Results
deep dive into prisma 2020 item 23 requires reporting results of all sensitivity analyses, including those where conclusions did not change. The SWiM guideline provides additional reporting recommendations for non-quantitative sensitivity analyses.
Best practices for reporting:
Table format. Present all sensitivity analyses in a single summary table with columns for: analysis description, number of studies included, pooled estimate with confidence interval, I-squared, and whether the conclusion changed
Forest plot overlay. For key sensitivity analyses, consider showing the restricted analysis alongside the primary analysis in a single forest plot
Narrative interpretation. State explicitly whether the primary conclusion was robust or sensitive to each analysis. Avoid burying important sensitivity results in supplementary materials
Example reporting language: "The primary analysis included 14 trials and found a pooled standardized mean difference of -0.45 (95% CI: -0.62 to -0.28) favoring the intervention. Restricting to the 8 trials with low risk of bias yielded a smaller but still significant effect (SMD -0.31, 95% CI: -0.52 to -0.10). Leave-one-out analysis showed that no single trial changed the direction or significance of the pooled estimate. Results were consistent when using a fixed-effect model (SMD -0.42, 95% CI: -0.55 to -0.29)."
Pre-Specifying Sensitivity Analyses in the Protocol
List each planned sensitivity analysis with justification. Example: "We will restrict the meta-analysis to studies rated as low risk of bias to assess whether pooled effects are driven by methodologically weaker studies."
Distinguish from subgroup analyses.Subgroup analyses explore effect modification by clinical characteristics. Sensitivity analyses test robustness to methodological choices. Some analyses could be either (e.g., restricting by study design), so label them clearly.
Limit the number. Running 20 sensitivity analyses inflates the chance of finding one that "works." Pre-specify 3-6 that address the most important decision nodes for your review question.
Allow for post hoc additions. State that additional sensitivity analyses may be conducted if unexpected issues arise during data extraction or analysis, and label these as exploratory.
Common Mistakes to Avoid
Running sensitivity analysis only when results are unexpected. If you only test robustness when the primary result surprises you, this introduces bias. Pre-specify analyses regardless of what you expect to find.
Interpreting a non-significant sensitivity analysis as "no effect." If restricting to low risk of bias studies makes the pooled effect non-significant, this does not prove the intervention is ineffective. It may simply reflect reduced statistical power from fewer studies. Report the point estimate and confidence interval, not just the p-value.
Dropping studies based on sensitivity results. Sensitivity analysis is diagnostic, not prescriptive. If leave-one-out reveals an influential study, investigate and discuss it. Do not remove it from the primary analysis without a pre-specified, methodologically justified reason.
Ignoring sensitivity analyses that change the conclusion. If one of your pre-specified analyses overturns the result, this is arguably the most important finding of your review. Reporting only the analyses that support your conclusion is a form of selective reporting.
Not running enough sensitivity analyses. A single leave-one-out analysis is better than nothing, but it only addresses one type of uncertainty. Aim for sensitivity analyses that cover study inclusion, risk of bias, statistical model, and at least one domain specific to your review question (e.g., missing data handling, dose categorization, follow-up duration).
Sensitivity analysis varies methodological decisions (inclusion criteria, model choice, risk of bias thresholds) to test whether conclusions change. Subgroup analysis splits the data by clinical or demographic characteristics (age groups, intervention type, setting) to test whether the effect differs across populations. Sensitivity analysis asks 'are results robust?' while subgroup analysis asks 'does the effect vary across groups?' Both should be pre-specified in the protocol.
Systematic reviews involve dozens of judgment calls: which studies to include, which data to extract when multiple time points exist, which model to use, and how to handle missing data. Sensitivity analysis reveals whether your conclusions depend on any single decision. If changing a reasonable methodological choice flips the conclusion, readers and guideline developers need to know the evidence is fragile on that point.
The most common approach is leave-one-out analysis: remove each study in turn, re-run the meta-analysis, and check whether the pooled estimate or its significance changes. Beyond leave-one-out, you can restrict analysis to low risk of bias studies only, change the statistical model (fixed vs. random effects), use different effect size measures, or exclude studies with imputed data. Pre-specify which sensitivity analyses you will run in your protocol.
Leave-one-out analysis sequentially removes one study at a time from the meta-analysis and recalculates the pooled estimate. If the pooled result remains stable regardless of which study is removed, findings are robust. If removing a single study changes the direction or significance of the result, that study is influential and should be investigated for clinical or methodological reasons that might explain its outsized impact.
Sensitivity analysis should be planned during the protocol stage and conducted after the primary analysis. At minimum, run a sensitivity analysis for study quality (restricting to low risk of bias studies), for model choice (fixed vs. random effects), and for any methodological decision where reasonable alternative choices existed. Additional sensitivity analyses may be triggered by unexpected heterogeneity or reviewer concerns during the analysis phase.
Share
Found this useful? Share it with your colleagues.
Need help with your meta-analysis?
Our PhD statisticians run complete meta-analyses: effect sizes, forest plots, heterogeneity testing, and publication-ready results sections.
Reading About Meta-Analysis? Our PhD Team Runs Them Every Day.
From data extraction to forest plots, sensitivity analysis, and a journal-ready manuscript. We handle the full meta-analysis so you can focus on your research question.
Our promise: Free re-run of the pooled analysis if reviewers question the estimate or model.
4.9 / 5Quote within a few hoursmetafor R + Cochrane HandbookPhD methodologistConfidential by default
Dr. Sarah Mitchell holds a PhD in Biostatistics from Johns Hopkins Bloomberg School of Public Health and has over 15 years of experience in systematic review methodology and meta-analysis. She has authored or co-authored 40+ peer-reviewed publications in journals including the Journal of Clinical Epidemiology, BMC Medical Research Methodology, and Research Synthesis Methods. A former Cochrane Review Group statistician and current editorial board member of Systematic Reviews, Dr. Mitchell has supervised 200+ evidence synthesis projects across clinical medicine, public health, and social sciences.
Reading About Meta-Analysis? Our PhD Team Runs Them Every Day.
From data extraction to forest plots, sensitivity analysis, and a journal-ready manuscript. We handle the full meta-analysis so you can focus on your research question.
A rigorous, doctoral-level guide to conducting a meta-analysis: defining the question, extracting effect sizes and their variances, choosing a between-study variance estimator, pooling, and diagnosing heterogeneity and bias.
Meta-analysis in psychology pools the effect sizes from many studies into one reliable result. Learn the definition, real examples, and how researchers run one.
Roughly 80 systematic reviews are published daily. The average takes 67.3 weeks, uses 5 authors, and costs about $141,195 in researcher time. Every figure sourced and linked.