latent.flows.model_garden_flow¶
model_garden_flow — Evaluate across multiple models with Pareto analysis.
Functions¶
model_garden_flow¶
model_garden_flow(eval_data: pd.DataFrame, judge_factory: Callable[[str], Any], models: list[str], gates: dict[str, float] | None = None, metric_directions: dict[str, bool] | None = None, score_types: dict[str, str] | None = None, rubrics: dict | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, pareto_x: str | None = None, pareto_y: str | None = None, save_plots: str | None = None) -> dict[str, Any]
Evaluate the same dataset across multiple models and produce Pareto analysis.
For each model, runs instrumented_judge_flow to get the full triad (accuracy + latency + tokens). Then computes Pareto frontier across models.
Args: eval_data: DataFrame with columns matching judge's prompt_template. judge_factory: Callable that takes a model name and returns a Judge instance. Example: lambda model: Judge("scorer", model=model, output_type=Scores) models: List of model identifiers to evaluate. gates: Optional quality thresholds (applied per-model). metric_directions: Optional metric direction overrides. score_types: Optional score type overrides. rubrics: Optional rubric overrides. n_resamples: Bootstrap resamples. confidence_level: CI confidence level. seed: Random seed. pareto_x: Metric name for Pareto x-axis (e.g. "latency_seconds"). pareto_y: Metric name for Pareto y-axis (e.g. "accuracy" or first score field). save_plots: Optional directory to save Pareto plot PNGs.
Returns: Dict with keys: - per_model: dict[model_name, instrumented_judge_flow result] - pareto: pareto_frontier result (pareto_optimal, dominated_by, table) - summary_table: list of dicts with model, metrics, is_pareto - figures: list of matplotlib figures (if pareto_x and pareto_y provided)