Skip to content

latent.stats.agreement

Multi-judge agreement and aggregation metrics (FR-2.4).

Provides inter-annotator agreement statistics (Cohen's kappa, Fleiss' kappa, Krippendorff's alpha) and aggregation utilities (majority vote, disagreement flagging) for multi-judge evaluation workflows.

Functions

cohens_kappa

cohens_kappa(judge_a: np.ndarray, judge_b: np.ndarray) -> float

Compute Cohen's kappa for pairwise agreement between two judges.

Cohen's kappa measures inter-rater reliability for two raters, adjusting for agreement expected by chance.

Args: judge_a: 1D array of category assignments from judge A. judge_b: 1D array of category assignments from judge B.

Returns: Cohen's kappa coefficient in [-1, 1]. 1.0 = perfect agreement, 0 = chance agreement, <0 = worse than chance.

Raises: ValueError: If arrays have different lengths.

disagreement_flags

disagreement_flags(judge_scores: np.ndarray, threshold: float = 0.5) -> np.ndarray

Flag items where judges disagree above a threshold.

Disagreement is computed as twice the minority vote proportion: 2 * min(mean, 1 - mean). This gives 0 when all judges agree and 1.0 for a 50/50 split. For example, a 3-1 split among 4 judges gives 2 * min(0.75, 0.25) = 0.50.

Args: judge_scores: 2D array of shape (n_items, n_judges) with binary values (0 or 1). threshold: Disagreement threshold in [0, 1]. Items with disagreement >= threshold are flagged (default 0.5).

Returns: Boolean array of shape (n_items,). True where disagreement is at or above the threshold.

fleiss_kappa

fleiss_kappa(ratings: np.ndarray) -> float

Compute Fleiss' kappa for multi-judge agreement.

Fleiss' kappa extends Cohen's kappa to any fixed number of raters, measuring the degree of agreement beyond what would be expected by chance.

Args: ratings: 2D array of shape (n_items, n_judges) with category assignments. Each cell contains the category label assigned by that judge to that item.

Returns: Fleiss' kappa coefficient. 1.0 = perfect agreement, 0 = chance agreement.

Raises: ValueError: If ratings is not 2D.

krippendorffs_alpha

krippendorffs_alpha(ratings: np.ndarray, level_of_measurement: str = 'ordinal') -> float

Compute Krippendorff's alpha for inter-rater reliability.

Supports different levels of measurement which affect the difference function used to compute disagreement.

Args: ratings: 2D array of shape (n_items, n_judges) with values. Missing values should be represented as np.nan. level_of_measurement: One of "nominal", "ordinal", "interval", "ratio".

Returns: Krippendorff's alpha in [-1, 1]. 1.0 = perfect reliability, 0 = absence of reliability.

Raises: ValueError: If level_of_measurement is not recognized.

majority_vote

majority_vote(judge_scores: np.ndarray) -> np.ndarray

Aggregate binary judge scores via majority vote.

For each item, returns 1 if more than half of judges scored 1, 0 otherwise. Ties (exactly 50%) resolve to 0.

Args: judge_scores: 2D array of shape (n_items, n_judges) with binary scores (0 or 1).

Returns: 1D array of consensus scores, shape (n_items,).