Back to Blog
Statistics
8 min read

Krippendorff's Alpha: Agreement Beyond Two Coders

Krippendorff's alpha explained: what it measures, how it differs from Cohen's kappa, acceptable thresholds, and which agreement coefficient fits your coding design.

Research Gold Team

September 22, 2026

Need the number rather than the theory? Our inter-rater agreement calculator computes Krippendorff's alpha alongside Cohen's and Fleiss' kappa, with a bootstrap confidence interval.

Key Takeaways

Krippendorff's alpha measures agreement among any number of coders, on any measurement level, with missing data allowed

It is computed as 1 minus the ratio of observed to expected disagreement, so 1 is perfect agreement and 0 is chance-level

Krippendorff's own convention is that alpha of 0.800 or above supports conclusions, and 0.667 to 0.800 supports tentative conclusions only

Cohen's kappa handles exactly two coders on nominal categories. Alpha generalises past both limits

Krippendorff's alpha and Cronbach's alpha are unrelated despite the shared name

Krippendorff's alpha is a reliability coefficient that measures how consistently independent coders assign the same values to the same units, after correcting for the agreement that chance alone would produce. Its distinguishing feature is generality: it accepts any number of coders, any measurement level from nominal to ratio, and missing data, because it is built from observed and expected disagreement rather than from a two-by-two agreement table. It is computed as one minus the ratio of observed disagreement to expected disagreement, which puts perfect agreement at 1 and chance-level agreement at 0.

That construction is why alpha turns up wherever real coding projects outgrow textbook conditions. Three coders instead of two, an ordinal severity code rather than a yes or no, one coder who did not get to the last twenty transcripts: each of those breaks Cohen's kappa and none of them breaks alpha.

Which agreement coefficient your design actually calls for

Choosing the coefficient is a design question, and the wrong choice either fails to run or quietly misreports. The decision comes down to three features of your coding setup:

  • Two coders, every unit coded, unordered categories. Cohen's kappa is the conventional and expected choice. It is the most widely recognised coefficient, and reviewers in health research look for it first.
  • Three or more coders, every unit coded, unordered categories. Fleiss' kappa extends the logic to multiple raters, treating them as interchangeable rather than as named individuals.
  • Any number of coders, ordered or continuous codes, or missing values. Krippendorff's alpha, with the difference function set to match the measurement level.
  • Continuous measurements where you care about absolute agreement rather than category matching. The intraclass correlation coefficient is the right family, and our intraclass correlation calculator covers the common forms.
  • Extreme prevalence, where one category dominates. Kappa can collapse towards zero even when coders agree on almost every unit, a behaviour known as the kappa paradox. Gwet's AC1 is more stable under those conditions and is worth reporting alongside.

Our inter-rater agreement calculator computes the whole set, so the practical route is to run the coefficient your design demands and report a second one as a robustness check rather than shopping for the highest number.

Need statistical analysis support?

Our PhD statisticians handle data analysis, produce reproducible R code, and write results sections that satisfy peer reviewers.

Setting the difference function correctly

Alpha needs to know how far apart two disagreeing codes are, and that is set by the metric, sometimes called the difference function. Getting it wrong is the most common technical mistake and it almost always costs you agreement you had actually achieved.

Nominal. Codes are unordered labels, so any disagreement counts fully. Use this for thematic codes, presence or absence, or category assignment where no code is closer to another.

Ordinal. Codes are ranked. A coder who said moderate where another said severe has disagreed less badly than one who said none. Use this for severity ratings, Likert-type codes, and confidence levels. Running nominal alpha on ordinal data treats every disagreement as total and systematically understates reliability.

Interval and ratio. Codes are measurements, and disagreement is proportional to numeric distance. Use these for counts, durations, and scores.

State the metric you used when you report the value. An alpha of 0.74 means different things under nominal and ordinal treatment of the same data, and a reader cannot interpret the number without knowing which was applied.

Reading the value honestly

Krippendorff's own convention, which is the one usually cited, is that data with alpha at or above 0.800 can support conclusions, that 0.667 to 0.800 supports tentative conclusions only, and that below 0.667 the coded data should not be relied on for conclusions. Some fields and some journals apply looser thresholds, particularly for exploratory coding of difficult constructs.

Two cautions matter more than the cut-off. First, alpha is an estimate, and on a modest number of units it is a noisy one. A value of 0.81 from 40 coded units is not meaningfully better than 0.78; reporting a bootstrap confidence interval alongside the point estimate is what tells a reader whether the threshold was cleared reliably. Second, alpha can be negative, which indicates systematic disagreement rather than merely poor agreement, and it is a signal that the codebook contains a definition two coders are reading in opposite directions.

Reporting a qualitative study? The COREQ checklist is where the coding and coder items are expected.

What a high alpha does not prove

This is the part most worth being direct about, because it is where agreement statistics get misused in qualitative and mixed-methods work. Alpha measures reliability, not validity. It tells you coders applied the same rules the same way. It cannot tell you the rules were worth applying.

Two coders trained on the same flawed codebook will agree with each other beautifully. An alpha of 0.92 on a coding frame that misses the phenomenon entirely is a precise measurement of nothing, and presenting it as evidence that the themes are sound is a category error that experienced qualitative reviewers pick up immediately. The credibility of a coding frame comes from how it was developed and checked against the data, which is a separate argument made in the methods. Our overview of qualitative evidence synthesis methods covers where that argument belongs.

The corollary is also true: a modest alpha is not automatically a failure. For genuinely interpretive constructs, agreement in the 0.70s with a documented reconciliation process can be more credible than a suspiciously high figure on codes so coarse that agreement was inevitable.

Where agreement statistics are expected in reporting

Coding reliability is not optional disclosure in most reporting guidelines, it is an item. The COREQ checklist for interview and focus group research asks how many coders coded the data and whether a coding tree was provided, and a reader will expect either an agreement statistic or a described reconciliation process. SRQR covers the same ground for broader qualitative designs. In systematic reviews, agreement between screeners is routinely reported at the title and abstract stage, and PRISMA 2020 expects the process for selecting studies to be described including how many reviewers worked independently.

For screening specifically, two coders on include or exclude decisions is the standard setup, which means kappa is the usual statistic and alpha is the fallback when a third screener resolved conflicts or when some records were only seen by one person.

Reporting it in a sentence

A complete report of alpha contains four things: the coefficient and its value, the metric used, the number of units and coders, and an interval. For example: agreement across three coders on 120 transcript segments was assessed with Krippendorff's alpha using an ordinal difference function, alpha equal to 0.83 with a 95 percent bootstrap confidence interval of 0.76 to 0.89.

That sentence is checkable, and it forecloses the three questions a methodological reviewer would otherwise ask. If you also ran a second coefficient as a robustness check, give it in the same place rather than choosing between them after seeing the results.

Pro Tip

Pick the metric before you run it

Alpha takes a difference function matched to your data: nominal for unordered categories, ordinal for ranked ones, interval or ratio for measurements. Running nominal alpha on ordinal codes throws away the information that near-misses are better than far ones, and usually understates agreement.

Pro Tip

Do not delete cases to make it run

The main practical reason to choose alpha is that it handles missing values directly. Dropping units that one coder skipped biases the estimate and defeats the point of using it.

Pro Tip

Report a confidence interval

Alpha on 40 units is a noisy estimate. A bootstrap interval tells a reader whether your value is reliably above the threshold or just happens to be this time.

Frequently Asked Questions

4
It measures the extent to which independent coders assign the same values to the same units, corrected for the agreement that would occur by chance. It is a reliability coefficient for coded data rather than a measure of validity, so a high alpha says coders applied a codebook consistently. It does not say the codebook captured anything worth measuring.
Cohen's kappa is defined for exactly two coders rating every unit on nominal categories. Krippendorff's alpha allows any number of coders, any measurement level including ordinal and interval data, and missing values, because it works from observed and expected disagreement rather than from a square agreement matrix. Where both apply, and the data are complete and nominal with two coders, they give very similar answers.
Krippendorff's own recommendation is to rely on data with alpha of 0.800 or above, and to draw only tentative conclusions where alpha falls between 0.667 and 0.800. Below 0.667 the data are generally considered unreliable for drawing conclusions. These are conventions rather than statistical rules, and some fields apply looser cut-offs, so state the threshold you adopted and cite its source.
Cronbach's alpha is a different coefficient and answers a different question, despite the shared surname. It estimates the internal consistency of a set of scale items measuring one construct, computed from the number of items and the ratio of the summed item variances to the total score variance. Krippendorff's alpha estimates agreement between coders. A questionnaire's reliability is Cronbach's territory; a codebook's reliability is Krippendorff's.
Share

Found this useful? Share it with your colleagues.

Need statistical analysis support?

Our PhD statisticians handle data analysis, produce reproducible R code, and write results sections that satisfy peer reviewers.

Explore our Biostatistics Service, handled end-to-end by a PhD methodologist.

Biostatistics Support

Need a Statistician? Our PhD Team Handles the Numbers.

From data cleaning to advanced statistical analysis, reproducible R code, and a results section ready for peer review. We handle the stats so you focus on the science.

Our promise: Free re-run and re-write if reviewers question the analysis or reporting.

4.9 / 5Quote within a few hoursReproducible R or Stata codePhD methodologistConfidential by default
Chat on WhatsApp now
RG

Written by

Research Gold Team

PhD-Level Methodologists

Our team comprises PhD-level methodologists with peer-reviewed publication records in systematic reviews, meta-analyses, and biostatistics. Every article is fact-checked against current Cochrane Handbook, PRISMA 2020, and JBI guidelines to ensure the advice you read reflects the latest evidence synthesis standards.

Research Gold handles coding, agreement statistics and the write-up. See our qualitative data analysis service or request a quote.

Need a Statistician? Our PhD Team Handles the Numbers.

From data cleaning to advanced statistical analysis, reproducible R code, and a results section ready for peer review. We handle the stats so you focus on the science.

Quote within a few hours. Pay only after you approve your quote. Unlimited revisions within your agreed scope. Confidential by default.