latent.stats.agreement¶
Multi-judge agreement and aggregation metrics (FR-2.4).
Provides inter-annotator agreement statistics (Cohen's kappa, Fleiss' kappa, Krippendorff's alpha) and aggregation utilities (majority vote, disagreement flagging) for multi-judge evaluation workflows.
Functions¶
cohens_kappa¶
Compute Cohen's kappa for pairwise agreement between two judges.
Cohen's kappa measures inter-rater reliability for two raters, adjusting for agreement expected by chance.
Args: judge_a: 1D array of category assignments from judge A. judge_b: 1D array of category assignments from judge B.
Returns: Cohen's kappa coefficient in [-1, 1]. 1.0 = perfect agreement, 0 = chance agreement, <0 = worse than chance.
Raises: ValueError: If arrays have different lengths.
disagreement_flags¶
Flag items where judges disagree above a threshold.
Disagreement is computed as twice the minority vote proportion:
2 * min(mean, 1 - mean). This gives 0 when all judges agree
and 1.0 for a 50/50 split. For example, a 3-1 split among 4
judges gives 2 * min(0.75, 0.25) = 0.50.
Args: judge_scores: 2D array of shape (n_items, n_judges) with binary values (0 or 1). threshold: Disagreement threshold in [0, 1]. Items with disagreement >= threshold are flagged (default 0.5).
Returns: Boolean array of shape (n_items,). True where disagreement is at or above the threshold.
fleiss_kappa¶
Compute Fleiss' kappa for multi-judge agreement.
Fleiss' kappa extends Cohen's kappa to any fixed number of raters, measuring the degree of agreement beyond what would be expected by chance.
Args: ratings: 2D array of shape (n_items, n_judges) with category assignments. Each cell contains the category label assigned by that judge to that item.
Returns: Fleiss' kappa coefficient. 1.0 = perfect agreement, 0 = chance agreement.
Raises: ValueError: If ratings is not 2D.
krippendorffs_alpha¶
Compute Krippendorff's alpha for inter-rater reliability.
Supports different levels of measurement which affect the difference function used to compute disagreement.
Args: ratings: 2D array of shape (n_items, n_judges) with values. Missing values should be represented as np.nan. level_of_measurement: One of "nominal", "ordinal", "interval", "ratio".
Returns: Krippendorff's alpha in [-1, 1]. 1.0 = perfect reliability, 0 = absence of reliability.
Raises: ValueError: If level_of_measurement is not recognized.
majority_vote¶
Aggregate binary judge scores via majority vote.
For each item, returns 1 if more than half of judges scored 1, 0 otherwise. Ties (exactly 50%) resolve to 0.
Args: judge_scores: 2D array of shape (n_items, n_judges) with binary scores (0 or 1).
Returns: 1D array of consensus scores, shape (n_items,).