latent.optimize.dspy.metrics¶
DSPy metric adapters and calibration metric.
Classes¶
BatchCalibrationMetric¶
Accumulating calibration metric for DSPy optimization.
Thread-safe. Accumulates (judge_label, human_label) pairs and
computes calibration stats once min_samples have been collected.
Returns 0.0 before the threshold is reached.
Functions¶
adapt_metric¶
Bridge a latent metric function to DSPy metric format.
Args: metric_fn: A callable (ground_truth, predicted) -> float. input_field: Field name on the dspy.Example for ground truth. output_field: Field name on the dspy.Prediction for the prediction.
Returns: A DSPy-compatible metric function (example, prediction, trace) -> float.
Raises:
KeyError: If output_field is missing from the prediction.
agreement_metric¶
Per-example agreement score: 1.0 if labels match, 0.0 otherwise.
For batch-level inter-rater reliability (Cohen's kappa), use BatchCalibrationMetric instead.
Methods¶
BatchCalibrationMetric.reset¶
Reset accumulated labels.