Back to Blog
Meta-Analysis
14 min read

Meta-Analysis in Psychology: Definition, Examples, and How It Works

Meta-analysis in psychology pools the effect sizes from many studies into one reliable result. Learn the definition, real examples, and how researchers run one.

Dr. James Whitfield

Running a psychology meta-analysis for your thesis or paper? Request a meta-analysis quote and we will scope the effect sizes, pooling model, and figures with you.

Key Takeaways

Meta-analysis in psychology statistically combines the effect sizes of multiple studies on the same question into a single pooled estimate that is more precise than any one study.

Gene Glass coined the term meta-analysis in 1976, and the Smith and Glass synthesis of psychotherapy outcome studies became the field's founding example.

Effect sizes such as Cohen's d and the correlation coefficient r are the common currency that lets results from different psychology studies be combined.

A random-effects model is usually preferred in psychology because the true effect realistically varies across populations, designs, and measures.

Heterogeneity (often summarized with I-squared) and publication bias are the two threats that every credible psychology meta-analysis must address.

Meta-analysis is a core response to the replication crisis because it weights evidence by precision rather than treating every single study as decisive.

Psychology runs two traditions: the Hedges-Olkin model that pools observed effect sizes, and Hunter-Schmidt psychometric meta-analysis that first corrects correlations for measurement unreliability and range restriction.

Meta-analysis in psychology is a statistical method that combines the effect sizes from many independent studies on the same question into a single pooled effect. Instead of treating any one experiment as the final word, it weights each study by its precision, so larger and more reliable studies carry more influence. The result is an estimate that is more accurate, and far harder to dismiss, than the scattered findings it draws from.

For a field built on small samples and surprising headlines, this matters enormously. A dozen studies on the same hypothesis can point in different directions purely by chance. Meta-analysis in psychology turns that noisy pile of results into one defensible number, complete with a confidence interval and an honest account of how much the studies disagree.

Why psychology leans so heavily on pooled evidence

Psychology has a structural problem that meta-analysis is unusually well suited to solve. Most individual studies recruit modest samples, which gives them low statistical power and wide margins of error. Two well-run experiments can easily produce opposite conclusions, not because the underlying truth changed, but because each caught a different slice of random variation.

This is the engine behind the replication crisis. When a single striking study fails to repeat, it is tempting to conclude the effect is fake. A better move is to ask what happens when you combine every credible study on the question. By pooling results and weighting them by precision, meta-analysis reveals whether an effect is real but small, genuinely absent, or strongly dependent on context. That shift, from celebrating one decisive study to weighing the whole body of evidence, is exactly why meta-analysis now sits at the top of the evidence hierarchy in psychological science.

The effect size: the common currency that makes pooling possible

You cannot average raw results from studies that used different scales, tasks, and sample sizes. What you can combine is the effect size, a standardized number that expresses the strength of a finding independently of the original measurement units. Two metrics dominate psychology. The first is Cohen's d, a standardized mean difference that reports how far two group means sit apart in standard-deviation units. The second is the correlation coefficient r, used when the question is about association rather than group difference.

The art of preparing a meta-analysis is converting every primary study into the same metric, then attaching a standard error to each one. Studies with tighter standard errors get more weight in the final weighted average. If you are still building intuition for these numbers, our guide to calculating standardized effect sizes walks through the conversions, and the effect size calculator handles the arithmetic for a single comparison.

Examples of meta-analysis in psychology

The method's founding example is also its clearest. In the mid-1970s Gene Glass coined the term meta-analysis, and together with Mary Lee Smith he synthesized hundreds of psychotherapy outcome studies. Their pooled result indicated that the average treated client improved more than roughly 75 percent of comparable untreated people, a conclusion no single trial could have established with the same authority. It reframed a contentious debate about whether therapy worked into a quantitative question about how much.

Other examples show the method's range. Cross-cultural syntheses of conformity experiments pooled decades of replications of the classic Asch line-judgement task to ask whether conformity has weakened over time and whether it differs across cultures. More recently, large coordinated efforts on ego depletion, the idea that self-control draws on a limited resource, found a much smaller effect than earlier optimistic summaries, precisely because they accounted for unpublished and underpowered studies. Each case demonstrates the same lesson: pooling many studies produces a more honest picture than any headline finding.

Need help with your meta-analysis?

Our PhD statisticians run complete meta-analyses: effect sizes, forest plots, heterogeneity testing, and publication-ready results sections.

Fixed-effect versus random-effects models

Once your effect sizes are ready, you choose how to pool them. A fixed-effect model assumes every study estimates one identical true effect, and any differences between them are pure sampling error. That assumption is rarely plausible in psychology, where samples, age groups, measures, and procedures vary from lab to lab.

The random-effects model is the realistic default. It assumes the true effect itself varies across studies around an average, so the analysis estimates both that average and the spread around it. In practice this gives smaller studies a fairer share of influence and produces wider, more honest confidence intervals. Choosing the model is not a cosmetic decision: it changes the pooled estimate, its precision, and how you interpret the result. The companion guide on how to do a meta-analysis details the between-study variance estimators (REML, Paule-Mandel) and the Hartung-Knapp adjustment that make a random-effects result defensible.

Two traditions: Hedges-Olkin and psychometric meta-analysis

Psychology is unusual in running two distinct meta-analytic frameworks, and a doctoral review should state which one it uses and why. The Hedges-Olkin tradition, described above and dominant in clinical and experimental psychology, pools observed effect sizes under fixed-effect or random-effects models and treats each study's effect as it was measured.

The Hunter-Schmidt tradition, also called psychometric meta-analysis or validity generalization, takes a different view that is influential in industrial-organizational, educational, and individual-differences research. It argues that an observed correlation is systematically biased downward by study artifacts, chiefly unreliable measurement and range restriction in the sample. Before pooling, it corrects each correlation for attenuation due to measurement error, dividing by the square root of the product of the two variables' reliabilities, and corrects for range restriction using the study's selection ratio. When reliabilities are reported inconsistently, it applies an artifact distribution to correct the pooled estimate as a whole rather than study by study. The payoff is an estimate of the true-score relationship between constructs, not merely the relationship between fallible measures, which is why personnel-selection research uses it to argue that a validity coefficient generalizes across settings once artifacts are removed.

The choice between the traditions is substantive. Psychometric corrections raise the pooled effect and widen its interval, and reviewers will expect you to justify the reliabilities and selection ratios you used, so commit to one framework in the protocol rather than switching after seeing the data.

Pooling effect sizes by hand invites mistakes. Build your forest plot from extracted study data in minutes.

Heterogeneity and publication bias: the two threats you must report

A pooled number is only trustworthy if you are candid about two problems. The first is heterogeneity, the genuine variation in the true effect from study to study. It is commonly summarized with the I-squared statistic, which expresses the percentage of variation that reflects real differences rather than chance. High heterogeneity is not a failure; it is a signal to investigate moderators such as age, design, or outcome measure. Our guide to measuring heterogeneity with I-squared explains how to interpret and act on it.

The second threat is publication bias, often called the file drawer problem. Studies with significant, exciting results get published; null results quietly disappear into desk drawers. If your meta-analysis only sees the published tip of the iceberg, it will overstate the effect. Funnel plots and formal tests help detect this asymmetry so you can flag it honestly.

Modern psychology goes well beyond the funnel plot, and a doctoral review is expected to use several of these tools. Egger's regression tests funnel asymmetry formally, trim-and-fill imputes the studies that asymmetry implies are missing, and PET-PEESE uses the relationship between effect size and standard error to estimate what an infinitely precise study would have found. p-curve and p-uniform analyze the distribution of statistically significant p-values to judge whether a literature carries genuine evidential value or merely reflects selective reporting and p-hacking, while three-parameter selection models estimate the effect under an explicit model of how significance shapes what gets published. The structural answer to the file-drawer problem is pre-registration and large multi-site Registered Replication Reports, such as the Many Labs projects, which publish results regardless of outcome and have repeatedly pulled once-celebrated effects toward zero.

The standard way to display all of this at once is the forest plot, which shows each study's effect size, its confidence interval, and the pooled diamond at the bottom. Learning to read one is a core skill, covered in our guide to reading a forest plot.

How a psychology meta-analysis comes together, step by step

The workflow is disciplined and reproducible. You define a precise question and pre-register the protocol. You search multiple databases and screen studies against explicit inclusion criteria. You extract the data needed to compute an effect size and its standard error from every included study. You choose a pooling model, run the synthesis, and produce a forest plot. Finally you probe heterogeneity, test for publication bias, and run sensitivity checks to see whether any single study is driving the result.

If you want the full procedure laid out in detail, with the formulas and software steps, see our companion guide on how to do a meta-analysis. When it is time to assemble your own figures, the forest plot generator turns extracted study data into a publication-ready plot.

Where meta-analysis fits in your research

A meta-analysis is the quantitative heart of a larger systematic review. The review is the transparent process of finding and appraising every relevant study; the meta-analysis is the statistical step that pools their numbers. Done well, the pair produces the most credible evidence psychology can offer on a question, which is why journals, funders, and thesis committees prize them.

That rigor is also why meta-analysis is demanding to execute. Extracting effect sizes consistently, choosing the right model, and defending every methodological decision takes time and statistical fluency. Whether you are a doctoral student synthesizing the literature for a thesis chapter or a research team preparing a publishable review, the difference between a defensible pooled estimate and a misleading one lives in those details.

Pro Tip

Pre-register your protocol

Register your question, inclusion criteria, and analysis plan before extraction. This prevents the selective decisions that quietly inflate effect sizes in psychology.

Pro Tip

Convert everything to one effect size metric

Decide early whether you will work in standardized mean difference or correlation, then convert all studies to that metric so they are genuinely comparable.

Pro Tip

Plan moderator analyses in advance

If you expect the effect to differ by age, design, or measure, specify those moderators upfront rather than fishing for them after seeing the data.

Pro Tip

Test for publication bias every time

Use funnel plots and formal tests so a tidy result is not just the visible tip of a larger, unpublished literature.

Frequently Asked Questions

8
A meta-analysis in psychology is a quantitative method that combines the effect sizes from several independent studies on the same question into one pooled estimate. Each study is weighted by its precision, so larger and more reliable studies count for more, producing a result that is more accurate than any single study on its own.
The classic example is Smith and Glass, who synthesized hundreds of psychotherapy outcome studies in the 1970s and reported that the average treated client improved more than roughly 75 percent of untreated people. Other well known examples include cross-cultural meta-analyses of conformity and large replication efforts on ego depletion.
Individual psychology studies often use small samples and produce conflicting results. Meta-analysis resolves that conflict by pooling effect sizes, increasing statistical power, and showing whether an effect is consistent. It has become central to the field's response to the replication crisis.
A systematic review is the structured process of finding, appraising, and summarizing all studies on a question. A meta-analysis is the optional statistical step within a systematic review that pools the numerical results. Every meta-analysis should sit inside a systematic review, but not every systematic review includes a meta-analysis.
There is no fixed minimum. A meta-analysis can technically be run on two studies, though estimates of heterogeneity become unstable with very few. Most psychology meta-analyses include at least ten studies so that moderator analyses and publication bias tests are meaningful.
Meta-analysis is quantitative. It works with numerical effect sizes and their standard errors. The qualitative counterpart, which synthesizes themes from non-numerical studies, is usually called a narrative synthesis or a qualitative evidence synthesis.
Psychometric meta-analysis, developed by Hunter and Schmidt and also called validity generalization, corrects each study's correlation for statistical artifacts before pooling, chiefly measurement unreliability (attenuation) and range restriction in the sample. The aim is to estimate the true-score relationship between constructs rather than the relationship between fallible measures. It is widely used in industrial-organizational and individual-differences research and typically yields a larger pooled effect with a wider interval than the Hedges-Olkin approach.
p-curve examines the distribution of statistically significant p-values across a set of studies. A literature with a real effect produces a right-skewed curve with many very small p-values, whereas a flat or left-skewed curve suggests the significant findings come from selective reporting or p-hacking rather than a true effect. It is used alongside PET-PEESE and selection models to assess whether a body of psychology research has genuine evidential value.
Share

Found this useful? Share it with your colleagues.

Need help with your meta-analysis?

Our PhD statisticians run complete meta-analyses: effect sizes, forest plots, heterogeneity testing, and publication-ready results sections.

Explore our Meta-Analysis Service, handled end-to-end by a PhD methodologist.

Meta-Analysis Support

Reading About Meta-Analysis? Our PhD Team Runs Them Every Day.

From data extraction to forest plots, sensitivity analysis, and a journal-ready manuscript. We handle the full meta-analysis so you can focus on your research question.

Our promise: Free re-run of the pooled analysis if reviewers question the estimate or model.

4.9 / 5Quote within a few hoursmetafor R + Cochrane HandbookPhD methodologistConfidential by default
Chat on WhatsApp now
Dr. James Whitfield

Written by

Dr. James Whitfield

Senior Reviewer
Risk of BiasGRADEClinical Trials

PhD in Clinical Trials Methodology. Owns risk-of-bias assessment using RoB 2, ROBINS-I V2, ROBINS-E, QUADAS-2, and the Newcastle-Ottawa Scale, plus GRADE certainty ratings on every review. Last methodological eye on every manuscript before release.

Ready to synthesize your psychology studies into a publishable result? Research Gold delivers full meta-analysis services, from effect size extraction to GRADE certainty ratings. Request your free consultation.

Reading About Meta-Analysis? Our PhD Team Runs Them Every Day.

From data extraction to forest plots, sensitivity analysis, and a journal-ready manuscript. We handle the full meta-analysis so you can focus on your research question.

Starting from the research question, not just the data? We run the whole systematic review and meta-analysis together. Quote my review + meta-analysis

Quote within a few hours. Pay only after you approve your quote. Unlimited revisions within your agreed scope. Confidential by default.