Skip to content

latent.optimize.dspy.metrics

DSPy metric adapters and calibration metric.

Classes

BatchCalibrationMetric

BatchCalibrationMetric(min_samples: int = 20, label_field: str = 'label')

Accumulating calibration metric for DSPy optimization.

Thread-safe. Accumulates (judge_label, human_label) pairs and computes calibration stats once min_samples have been collected. Returns 0.0 before the threshold is reached.

Functions

adapt_metric

adapt_metric(metric_fn: Callable, input_field: str, output_field: str) -> Callable

Bridge a latent metric function to DSPy metric format.

Args: metric_fn: A callable (ground_truth, predicted) -> float. input_field: Field name on the dspy.Example for ground truth. output_field: Field name on the dspy.Prediction for the prediction.

Returns: A DSPy-compatible metric function (example, prediction, trace) -> float.

Raises: KeyError: If output_field is missing from the prediction.

agreement_metric

agreement_metric(example: Any, prediction: Any, trace: Any = None) -> float

Per-example agreement score: 1.0 if labels match, 0.0 otherwise.

For batch-level inter-rater reliability (Cohen's kappa), use BatchCalibrationMetric instead.

Methods

BatchCalibrationMetric.reset

reset() -> None

Reset accumulated labels.