latent.stats.calibration¶
Calibration set management for judge evaluation (FR-2.1).
Computes sensitivity (TPR) and specificity (1-FPR) from calibration data where human labels serve as ground truth for evaluating judge accuracy.
Reference: Delegates to sklearn.metrics.recall_score.
Functions¶
compute_calibration¶
Compute sensitivity (TPR) and specificity (1-FPR) from calibration data.
For binary metrics, computes direct TPR/FPR from a confusion-matrix comparison of judge predictions vs. human ground-truth labels.
Args: judge_labels: 1D binary array of judge predictions (0 or 1). human_labels: 1D binary array of human ground-truth labels (0 or 1).
Returns: CalibrationStats with sensitivity, specificity, and sample size.
Raises: ValueError: If arrays have different lengths or n < 20.