Qualitative Analysis/Inter-rater reliability
Compare Coding
How to calculate inter-rater reliability in qualitative research
When two researchers code the same transcripts independently, how do you know whether they are applying codes in the same way? This guide explains Cohen's Kappa, how to calculate it in AnalyZ Solutions, how to interpret the result, and what to do when agreement is not yet good enough.
8 min read
Compare Coding
Intermediate
Inter-rater reliability (IRR) is a measure of how consistently two independent researchers apply the same coding scheme to the same data. It is one of the most important quality checks in qualitative analysis, particularly for studies that will be peer-reviewed, published, or used to inform policy decisions.
01Why inter-rater reliability matters
Qualitative coding involves interpretation. Two trained researchers reading the same passage may legitimately disagree about which code applies, particularly for codes with fuzzy boundaries or overlapping definitions. That disagreement is not necessarily a problem: it is often the starting point for a productive conversation about what the code really means.
The problem arises when disagreement is systematic and unexamined. If Coder A applies "Perceived barrier" to any mention of difficulty, while Coder B applies it only to explicit statements of inability, their coded datasets are not comparable. Any frequency analysis, comparative analysis, or thematic claim built on both datasets will be unreliable.
Measuring IRR forces the issue. It tells you which codes are being applied consistently and which are not, so you can resolve disagreements before analysis proceeds.
IRR is a process check, not a final score. A Kappa of 0.45 does not mean your analysis is poor, it means specific codes need sharper definitions. The value of computing Kappa is that it tells you exactly which codes to focus on, not that it passes or fails the study as a whole.
02What Cohen's Kappa measures
Cohen's Kappa (κ) measures agreement between two coders, corrected for the agreement that would occur by chance alone. If two coders randomly assigned codes, some passages would receive the same code by accident. Kappa subtracts this chance agreement from the observed agreement, giving a more conservative measure of true consistency.
The formula is: κ = (observed agreement − chance agreement) / (1 − chance agreement)
A Kappa of 1.0 means perfect agreement. A Kappa of 0 means the coders agree only as often as chance would predict. A negative Kappa means agreement is worse than chance, which almost always indicates a misunderstanding of the coding task.
| Kappa value | Interpretation | Action |
| 0.80 and above | Almost perfect agreement | Proceed to analysis with confidence |
| 0.60 to 0.79 | Substantial agreement | Acceptable: note any borderline codes and discuss |
| 0.40 to 0.59 | Moderate agreement | Revise code definitions, recalibrate, recode |
| 0.20 to 0.39 | Fair agreement | Significant revision needed before analysis |
| Below 0.20 | Slight or no agreement | Code definitions unclear: rethink and restart |
The widely accepted minimum for qualitative research reporting is κ ≥ 0.60. Studies published in peer-reviewed journals typically report κ ≥ 0.70 or above.
03How to calculate Kappa in AnalyZ Solutions
- Each coder completes their independent codingBoth coders code the same set of transcripts independently, without discussing their decisions. This independence is essential: if coders consult each other during coding, the reliability calculation is meaningless.
- Each coder exports their coding as JSONIn Thematic Coding, go to Review Codes and click the JSON (IRR) export button. This downloads a structured file containing all coded segments, the codes applied to each, and the codebook. Both coders export their own file separately.
- Open Compare Coding from qualitative moduleCompare Coding is a standalone module accessible directly from the qualitative module cards. It does not require transcripts to be loaded. Make sure both coders' files are available in the same computer.
- Load both export filesUpload each coder's JSON file and give each a label (e.g. the coder's name or initials). The labels appear in the results table.
- Click Compute agreement scoreAnalyZ Solutions compares the two files segment by segment, calculates Cohen's Kappa for each code individually, and computes a macro-average Kappa across all codes.
Watch a quick demonstration of comparing coding done by two coders in AnalyZ Solutions
04Reading the results
The results screen shows:
- Overall Kappa (macro-average): the average Kappa across all codes. This is the headline figure, but individual code Kappas are usually more informative.
- Per-code Kappa: a bar showing each code's Kappa and its strength label. This immediately identifies which codes are driving any disagreement.
- Percentage agreement: the raw proportion of segments where both coders agreed, shown alongside Kappa. Useful for context: a code with only 3 segments may show 100% agreement but the Kappa is not meaningful at that sample size.
The interpretation panel below the results classifies each code as strong (κ ≥ 0.80), acceptable (0.60–0.79), or below threshold (κ < 0.60), and provides specific next steps for each category.
05What to do when agreement is low
Low Kappa for a specific code almost always traces back to one of three causes: the code definition is ambiguous, the coders understood the definition differently, or the code is genuinely difficult to apply consistently to this type of data.
- Compare the two coders' segments side by side
Export both files and look at which passages one coder coded that the other did not. Look for a pattern: is one coder coding more broadly? More narrowly? Using the code for a different type of passage?
- Discuss the disagreements
Bring the disagreeing passages to a calibration session. For each passage, each coder explains why they made the decision they did. These conversations almost always surface the implicit difference in understanding.
- Revise the code definition
Update the code description and analytical memo to reflect the agreed definition. Add concrete examples of what qualifies and what does not. Make the boundary cases explicit.
- Recode a subset of transcripts
After calibration, have both coders independently recode two or three transcripts using the revised definition. Re-export and re-compute Kappa to confirm improvement before recoding the full dataset.
Do not average away low-Kappa codes. A high overall Kappa with one code at 0.30 should not be reported as acceptable. The low-Kappa code must be addressed, either resolved through calibration, or removed from the codebook if it cannot be applied reliably. Any findings built on that code are not trustworthy until agreement is established.
06When to use IRR and when it is not required
IRR is most important in studies where:
- Two or more researchers are contributing to the coding and analysis
- The findings will be peer-reviewed or published
- The coding framework will be applied to additional data in the future
- The study makes strong causal or prevalence claims based on coded data
Single-researcher studies, which are common in applied evaluation, can report rigour through other means: transparent codebook documentation, analytical memos, member checking, and reflexivity statements. IRR is a tool for demonstrating consistency between coders, not the only marker of rigour in qualitative research.
Frequently Asked Questions
Do both coders need to code all transcripts, or only a subset?
For a full reliability study, both coders code all transcripts independently. In practice, this is often resource-intensive. A common compromise is to have both coders code a random subset (typically 20 to 30 percent of transcripts) and calculate Kappa on that subset. The remaining transcripts are then coded by one coder, with the calibrated codebook applied consistently.
What if two codes consistently co-occur in one coder's data but not the other's?
This usually indicates that one coder is splitting a theme into two codes that the other is treating as one. Review whether the two codes are genuinely distinct or whether they should be merged. If they are kept separate, make the distinction explicit in both code definitions with concrete examples of each.
Can I report percentage agreement instead of Kappa?
Percentage agreement alone is not sufficient for peer-reviewed reporting because it does not correct for chance. Two coders who randomly applied codes would still agree some proportion of the time. Kappa corrects for this and is the expected standard. Report both Kappa and percentage agreement: the Kappa for rigour, the percentage for interpretability.
My two coders have very different levels of experience. Does that affect Kappa?
It can. A less experienced coder may apply codes more broadly or inconsistently, reducing Kappa. If possible, train both coders on the codebook together using example passages before independent coding begins. Calibration sessions, where coders practice on a few transcripts together and discuss disagreements, substantially improve agreement before the formal reliability coding starts.
What does it mean if Kappa is high but the two coders coded very different numbers of segments?
This suggests one coder is being more selective about what qualifies as a codeable segment. High Kappa with large differences in segment counts can indicate that one coder is missing relevant passages. Review the passages one coder coded that the other did not: these are the segments where the selection threshold differs most.
A code shows high percentage agreement but Kappa of 0.00. What does that mean?
This is one of the most important distinctions in reliability analysis, and it is easy to misread. High percentage agreement with Kappa near zero almost always means the code is rare in the data. If "Future direction", for example, appears in only 3 of 77 segments, both coders will say "no" to the other 74 automatically. Those 74 shared negatives inflate the percentage agreement to 96% even if the two coders never agreed on a single positive case. Kappa strips out this chance agreement and reveals that the coders are not actually applying the code consistently. Percentage agreement is misleading for rare codes: only Kappa is informative. The practical response is to look specifically at the excerpts where at least one coder applied the code and determine whether the definition needs clarification.
Is it acceptable to remove a code with low Kappa from the analysis?
It can be, but only after a genuine attempt to resolve the disagreement through calibration. If two coders cannot reach consistent application of a code after discussing the definition and recoding a subset of transcripts, the code may be too ambiguous to use reliably. Removing it and noting this as a limitation is more defensible than retaining it and building findings on unreliable coding. Do not remove a code simply because it is inconvenient: if it captures something theoretically significant, invest the time to sharpen the definition before giving up on it.
How do I report Kappa in a research report or paper?
Report the overall macro-average Kappa alongside the per-code Kappas, especially for any codes central to the main findings. A typical reporting sentence: "Inter-rater reliability was assessed using Cohen's Kappa on a random subset of 30% of transcripts coded independently by both researchers. Overall Kappa was 0.82 (almost perfect agreement). Kappa for individual codes ranged from 0.71 to 0.95, with the exception of one code (Future direction, κ = 0.43) which was revised after a calibration session; post-revision Kappa for that code was 0.78." Always note whether the Kappa reported is pre- or post-calibration, and whether it was computed on all transcripts or a subset.
What is the difference between Cohen's Kappa and weighted Kappa?
Cohen's Kappa treats all disagreements as equally serious: a coder applying code A when another applies code B counts the same as any other disagreement. Weighted Kappa allows some disagreements to be treated as less serious than others, which makes sense when codes are ordered or hierarchical (for example, rating scales where disagreeing by one step is less serious than disagreeing by three). For most qualitative coding tasks, where codes are nominal categories without an inherent order, standard Cohen's Kappa is the appropriate measure. Weighted Kappa is more relevant for rating or scoring tasks than for thematic coding.
We reached high Kappa on a subset. Do we need to verify on the full dataset?
Not necessarily, but there are circumstances where spot-checking later transcripts is worthwhile. If data collection covered multiple sites, time periods, or participant types, and the reliability subset came only from one context, coders may drift when they encounter different material. A brief check after coding is complete, comparing coding decisions on two or three additional transcripts from other contexts, costs little and can catch late-stage drift before it affects the final analysis.
Compare your coding in AnalyZ Solutions
Export your coding from Thematic Coding and calculate Kappa in minutes.
Get started
Related Guides