Krippendorff's alpha is a reliability coefficient that measures how consistently independent coders assign the same values to the same units, after correcting for the agreement that chance alone would produce. Its distinguishing feature is generality: it accepts any number of coders, any measurement level from nominal to ratio, and missing data, because it is built from observed and expected disagreement rather than from a two-by-two agreement table. It is computed as one minus the ratio of observed disagreement to expected disagreement, which puts perfect agreement at 1 and chance-level agreement at 0.
That construction is why alpha turns up wherever real coding projects outgrow textbook conditions. Three coders instead of two, an ordinal severity code rather than a yes or no, one coder who did not get to the last twenty transcripts: each of those breaks Cohen's kappa and none of them breaks alpha.
Choosing the coefficient is a design question, and the wrong choice either fails to run or quietly misreports. The decision comes down to three features of your coding setup:
- Two coders, every unit coded, unordered categories. Cohen's kappa is the conventional and expected choice. It is the most widely recognised coefficient, and reviewers in health research look for it first.
- Three or more coders, every unit coded, unordered categories. Fleiss' kappa extends the logic to multiple raters, treating them as interchangeable rather than as named individuals.
- Any number of coders, ordered or continuous codes, or missing values. Krippendorff's alpha, with the difference function set to match the measurement level.
- Continuous measurements where you care about absolute agreement rather than category matching. The intraclass correlation coefficient is the right family, and our intraclass correlation calculator covers the common forms.
- Extreme prevalence, where one category dominates. Kappa can collapse towards zero even when coders agree on almost every unit, a behaviour known as the kappa paradox. Gwet's AC1 is more stable under those conditions and is worth reporting alongside.
Our inter-rater agreement calculator computes the whole set, so the practical route is to run the coefficient your design demands and report a second one as a robustness check rather than shopping for the highest number.
Alpha needs to know how far apart two disagreeing codes are, and that is set by the metric, sometimes called the difference function. Getting it wrong is the most common technical mistake and it almost always costs you agreement you had actually achieved.
Nominal. Codes are unordered labels, so any disagreement counts fully. Use this for thematic codes, presence or absence, or category assignment where no code is closer to another.
Ordinal. Codes are ranked. A coder who said moderate where another said severe has disagreed less badly than one who said none. Use this for severity ratings, Likert-type codes, and confidence levels. Running nominal alpha on ordinal data treats every disagreement as total and systematically understates reliability.
Interval and ratio. Codes are measurements, and disagreement is proportional to numeric distance. Use these for counts, durations, and scores.
State the metric you used when you report the value. An alpha of 0.74 means different things under nominal and ordinal treatment of the same data, and a reader cannot interpret the number without knowing which was applied.
Krippendorff's own convention, which is the one usually cited, is that data with alpha at or above 0.800 can support conclusions, that 0.667 to 0.800 supports tentative conclusions only, and that below 0.667 the coded data should not be relied on for conclusions. Some fields and some journals apply looser thresholds, particularly for exploratory coding of difficult constructs.
Two cautions matter more than the cut-off. First, alpha is an estimate, and on a modest number of units it is a noisy one. A value of 0.81 from 40 coded units is not meaningfully better than 0.78; reporting a bootstrap confidence interval alongside the point estimate is what tells a reader whether the threshold was cleared reliably. Second, alpha can be negative, which indicates systematic disagreement rather than merely poor agreement, and it is a signal that the codebook contains a definition two coders are reading in opposite directions.
This is the part most worth being direct about, because it is where agreement statistics get misused in qualitative and mixed-methods work. Alpha measures reliability, not validity. It tells you coders applied the same rules the same way. It cannot tell you the rules were worth applying.
Two coders trained on the same flawed codebook will agree with each other beautifully. An alpha of 0.92 on a coding frame that misses the phenomenon entirely is a precise measurement of nothing, and presenting it as evidence that the themes are sound is a category error that experienced qualitative reviewers pick up immediately. The credibility of a coding frame comes from how it was developed and checked against the data, which is a separate argument made in the methods. Our overview of qualitative evidence synthesis methods covers where that argument belongs.
The corollary is also true: a modest alpha is not automatically a failure. For genuinely interpretive constructs, agreement in the 0.70s with a documented reconciliation process can be more credible than a suspiciously high figure on codes so coarse that agreement was inevitable.
Coding reliability is not optional disclosure in most reporting guidelines, it is an item. The COREQ checklist for interview and focus group research asks how many coders coded the data and whether a coding tree was provided, and a reader will expect either an agreement statistic or a described reconciliation process. SRQR covers the same ground for broader qualitative designs. In systematic reviews, agreement between screeners is routinely reported at the title and abstract stage, and PRISMA 2020 expects the process for selecting studies to be described including how many reviewers worked independently.
For screening specifically, two coders on include or exclude decisions is the standard setup, which means kappa is the usual statistic and alpha is the fallback when a third screener resolved conflicts or when some records were only seen by one person.
A complete report of alpha contains four things: the coefficient and its value, the metric used, the number of units and coders, and an interval. For example: agreement across three coders on 120 transcript segments was assessed with Krippendorff's alpha using an ordinal difference function, alpha equal to 0.83 with a 95 percent bootstrap confidence interval of 0.76 to 0.89.
That sentence is checkable, and it forecloses the three questions a methodological reviewer would otherwise ask. If you also ran a second coefficient as a robustness check, give it in the same place rather than choosing between them after seeing the results.