Cohen's Kappa Calculator

Measure inter-rater agreement corrected for chance.

Observed agreement: 80%

0.6

Expected (chance) agreement: 50%

How Cohen's kappa works

Cohen's kappa measures agreement between two raters classifying the same items into categories, correcting for the agreement you'd expect by chance alone — a more honest measure of true agreement than a raw percentage.

Enter the counts for all four combinations: both raters said yes, both said no, and the two disagreement combinations. The calculator returns observed agreement, expected (chance) agreement, and kappa.

This version handles the common two-rater, binary-classification case (yes/no, pass/fail, agree/disagree). For more than two raters or more than two categories, a generalized kappa or a different agreement statistic is needed.

The formula

κ = (Observed agreement − Expected agreement) ÷ (1 − Expected agreement)

Observed agreement is the share of items both raters classified the same way. Expected agreement is what you'd expect by chance, computed from each rater's marginal "yes" rate. Kappa scales the actual improvement over chance to the maximum possible improvement.

Worked example

100 items: 40 both-yes, 40 both-no, 10 + 10 disagreements

Observed agreement = 80%. Each rater says "yes" 50% of the time, giving expected chance agreement of 50%. κ = (0.80 − 0.50) ÷ (1 − 0.50) = 0.60 — commonly described as "substantial" agreement.

Frequently asked questions

What counts as a "good" kappa value?

A commonly cited scale (Landis & Koch, 1977) treats 0.61–0.80 as substantial agreement and above 0.80 as almost perfect — but interpretation guidelines vary by field, so check what's standard in your discipline.

Can kappa be negative?

Yes — a negative kappa means agreement is worse than chance, which usually signals a systematic disagreement or a coding-definition problem between raters worth investigating directly.

Why not just report percent agreement?

Raw percent agreement doesn't account for agreement that would happen by chance alone, especially when one category is much more common than the other — kappa corrects for that.

Next steps in your analysis

Built by PanelRoster

PanelRoster is an enterprise operations management platform for organizations coordinating distributed teams, assignments, workflows, quality assurance and operational execution.

Explore PanelRoster