Health & public health research support · led by a published academiccontact@researchflo.coReplies within 1 business day
Home / Blog / Reliability

When you need Cohen's kappa — and when you don't

Inter-coder reliability is expected in some traditions and actively inappropriate in others. How to tell which one you are in.

Reliability21 July 20266 min read
A laptop screen showing interview text with colour-coded highlighted passages

Reviewers ask for inter-coder reliability more often than they should, and researchers report it more often than they can justify. Both problems come from the same place: treating a statistic that belongs to one epistemological tradition as a general marker of quality.

When it genuinely applies

Kappa answers a specific question: if two people apply this coding scheme independently, how much do they agree beyond what chance would produce? That question only makes sense if there is a correct answer to agree on — a coding frame fixed in advance, applied to bounded units, with codes intended to be mutually exclusive.

  • Quantitative content analysis where codes will be counted and compared statistically
  • Deductive coding against a pre-existing framework or established instrument
  • Team coding where several people code different portions of one dataset
  • Clinical, systematic-review and health-services work where the field expects it

When it does not

In reflexive thematic analysis, Braun and Clarke are explicit: the researcher's interpretive position is a resource, not a source of error. Two analysts producing different readings is not a defect to be measured away — it is the thing the method assumes. Reporting kappa there does not strengthen the paper; it signals that the method has been misunderstood.

The same applies to most grounded theory, phenomenological and narrative work. Rigour in those traditions is demonstrated through the audit trail, reflexivity, negative case analysis and thick description — not through an agreement coefficient.

If a reviewer asks for kappa on a reflexive thematic analysis, the right response is a paragraph explaining why, not a statistic.

Reading the number

When kappa does apply, interpret it with the base rates in mind. The conventional bands — above 0.80 strong, 0.61 to 0.80 substantial, 0.41 to 0.60 moderate — are rules of thumb, and they mislead badly when a code is rare. A code appearing in 3% of segments can produce a low kappa alongside 96% raw agreement, because chance agreement is already very high. Report both the coefficient and the observed agreement, and say how prevalent the code was.

Use Cohen's kappa for two coders, Fleiss' kappa for three or more, and Krippendorff's alpha when you have missing data or ordinal categories. Double-code a meaningful sample — 10% to 25% of the dataset is the usual expectation, selected at random, not the easy transcripts.

What to do with disagreement

The number is diagnostic, not decorative. Every disagreement is a place where the codebook is ambiguous. Go through them, work out which definition was unclear, fix it, and record the fix. A project that reports kappa without describing how disagreements were resolved has reported half a result.

A quiet university library study space in warm daylight
Next step

Send us the brief. Get a price.

One email with your data type, volume and deadline is enough for a fixed written quote — back with you within 1 business day, NDA first if you prefer.

Chat on WhatsApp