Skip to content

latent.stats.calibration

Calibration set management for judge evaluation (FR-2.1).

Computes sensitivity (TPR) and specificity (1-FPR) from calibration data where human labels serve as ground truth for evaluating judge accuracy.

Reference: Delegates to sklearn.metrics.recall_score.

Functions

compute_calibration

compute_calibration(judge_labels: np.ndarray, human_labels: np.ndarray) -> CalibrationStats

Compute sensitivity (TPR) and specificity (1-FPR) from calibration data.

For binary metrics, computes direct TPR/FPR from a confusion-matrix comparison of judge predictions vs. human ground-truth labels.

Args: judge_labels: 1D binary array of judge predictions (0 or 1). human_labels: 1D binary array of human ground-truth labels (0 or 1).

Returns: CalibrationStats with sensitivity, specificity, and sample size.

Raises: ValueError: If arrays have different lengths or n < 20.