Skip to content

latent.flows.agent_garden_flow

agent_garden_flow — Run agent_eval_flow against multiple agents and compare.

Mirrors the role model_garden_flow plays for judges, but for the agent under test. For each agent_model in models, runs agent_eval_flow with that agent and a shared judge, then merges the per-variant reports into a single report_mode="comparison" StatisticalReport with:

  • variants — one VariantInfo per agent_model
  • records — every variant's records, with RecordResult.tag set to the full model identifier (e.g. "gemini/gemini-2.5-flash", not the short label) so per-row provenance is unambiguous
  • metrics — every variant's metrics, renamed <short-label>:<metric> and tagged with the full model in variant
  • gates — every variant's gates, metric_name renamed the same way
  • comparisons — every non-baseline variant against the baseline, one per judge score, family-wise corrected
  • power_analyses — one per binary-score comparison only; post_hoc_power's MDE is a two-proportion quantity, so it shares units with the observed effect nowhere else
  • category_summaries — every variant's, metrics renamed and tagged
  • failure_modes — every variant's, failure_mode renamed the same way (FailureModeSummary carries no variant field)

<short-label>:<metric> is one spelling shared by metrics and gates, and that is what makes the merged report gateable: publish_report and thresholds.yaml re-gating both index metrics by name and refuse to pick a winner among duplicates, so bare per-variant metric names would leave the report ungateable from either side.

baseline= names the reference arm; it defaults to the first model in models (is_baseline=True).

Typical use:

from latent.flows.agent_garden_flow import agent_garden_flow

result = agent_garden_flow(
    eval_data,
    agent_factory_for_model=lambda m: lambda: make_agent(model=m),
    judge=make_judge(),
    models=["gemini/gemini-2.5-flash", "gemini/gemini-2.5-flash-lite"],
    gates={"gemini-2.5-flash:accuracy": 0.6},
    seed=42,
)

report = result["report"]            # report_mode="comparison"
per_model = result["per_model"]      # dict[str, agent_eval_flow result]

Functions

agent_garden_flow

agent_garden_flow(eval_data: pd.DataFrame, agent_factory_for_model: Callable[[str], Callable[[], Any]], judge: Any, models: list[str], baseline: str | None = None, question_column: str = 'question', expected_column: str | None = None, category_column: str | None = None, id_column: str | None = None, context_columns: list[str] | None = None, context_formatter: Callable[[dict[str, Any]], str] | None = None, gates: dict[str, float | GateSpec] | None = None, scorers: list[Callable[[pd.DataFrame, list[dict]], dict[str, MetricResult]]] | None = None, failure_taxonomy: dict[str, str] | None = None, failure_filter_fn: Callable | None = None, failure_timeout_s: float | None = None, failure_concurrency: int | None = None, concurrency: int = 1, judge_concurrency: int | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, adjust_method: str | None = 'holm', summary_model: str | None = None, summary_label: str | None = None) -> dict[str, Any]

Run agent_eval_flow per agent_model and merge into one comparison report.

Args: eval_data: DataFrame with at least a question column. agent_factory_for_model: Callable that takes a model name and returns an agent factory (a zero-arg callable that produces a fresh agent instance configured for that model). The double-callable shape mirrors agent_eval_flow's agent_factory arg. judge: A shared Judge — or list of Judges — applied to every variant's outputs, exactly as agent_eval_flow takes it. models: List of agent model identifiers to compare, each listed once. Two models that shorten to the same label (openai/gpt-4o and azure/gpt-4o) both keep their full identifier as the label. baseline: The model every other variant is compared against. Defaults to models[0]. Naming it explicitly is worth doing: the baseline fixes the sign and the interpretation text of every delta, so reordering models in a params file silently inverts the whole report. question_column / expected_column / category_column / id_column: Forwarded to agent_eval_flow. context_columns / context_formatter: Forwarded to agent_eval_flow. gates: Quality thresholds. A <label>:<metric> key gates only that variant — "gemini-2.5-flash:accuracy" is a floor on that model alone. Passed programmatically, a bare key ("accuracy") gates every variant; from a thresholds.yaml section a bare key is rejected, because gates and metrics come out of the merge named <label>:<metric> and the teardown re-gates the merged report by name. scorers: Forwarded to agent_eval_flow per variant — the same deterministic scorers run against every model. failure_taxonomy / failure_filter_fn / failure_timeout_s / failure_concurrency / concurrency / judge_concurrency / n_resamples / confidence_level / seed: Forwarded to agent_eval_flow. adjust_method: Family-wise correction over the comparison p-values — one family spanning every (score, contender) pair, since they all come off the same eval set. "holm" (default) controls FWER, "fdr_bh" controls FDR, None leaves them uncorrected. summary_model: Optional model identifier for client-side summary generation. If set, generates a 3-5 sentence summary of the combined report and attaches it to report.summary. summary_label: Optional label passed to the summarizer. For a garden run this should reflect that it's a comparison — e.g. "Docs QA — gemini-flash vs gemini-flash-lite" — so the summary text frames metrics in comparison terms rather than as a single-agent baseline.

Returns: Dict with: - per_model: {model_name: agent_eval_flow result dict} - report: combined StatisticalReport with report_mode="comparison" - markdown: rendered markdown of the combined report - all_passed: True iff every variant's gates all passed

Raises: ValueError: models is empty or repeats a model, baseline is not among them, id_column repeats an id, or a gate key names an unknown variant.

Attributes

QUALIFIER

Separator between a variant label and a metric name.

A prefix, and this character, for two reasons measured rather than assumed: MLflow validates metric names against ^[/\w.\- :]*$, so @ makes log_metric raise and log_to_mlflow's catch-all swallows it — the whole merged report writes no metrics, params or artifacts. And gates.slices decodes a slice as {anchor}_{...}, a prefix relation, so a suffix qualifier would stop every scores_by slice being exempt from threshold coverage.