Skip to content

latent.flows.instrumented_flow

instrumented_flow — Score with triad reporting: accuracy + latency + tokens.

Functions

instrumented_judge_flow

instrumented_judge_flow(eval_data: pd.DataFrame, judge: Any, gates: dict[str, float] | None = None, metric_directions: dict[str, bool] | None = None, score_types: dict[str, str] | None = None, rubrics: dict | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> dict[str, Any]

Score a dataset with triad reporting: accuracy + latency + tokens.

Like judge_flow, but additionally captures per-call latency and token usage, reporting them alongside accuracy metrics with proper CIs.

Latency and token metrics emit mean (via the analyze step) plus p50/p90/p99 quantiles (quantile bootstrap). All latency/token metrics are automatically marked as lower_is_better for correct gating.

Args: eval_data: DataFrame with columns matching judge's prompt_template. judge: A Judge[T] instance with evaluate(row) -> T. gates: Optional dict mapping metric name to threshold. metric_directions: Optional dict mapping metric name to lower_is_better. Latency and token metrics are automatically added as lower_is_better=True. score_types: Optional override for score type detection. rubrics: Optional override for rubric detection. n_resamples: Bootstrap resamples for analysis. confidence_level: CI confidence level. seed: Random seed for reproducibility.

Returns: Dict with keys: scored_data, report, markdown, all_passed, call_metrics. call_metrics is a list of CallMetrics objects per row.