latent.flows.agent_garden_flow¶
agent_garden_flow — Run agent_eval_flow against multiple agents and compare.
Mirrors the role model_garden_flow plays for judges, but for the agent
under test. For each agent_model in models, runs agent_eval_flow with
that agent and a shared judge, then merges the per-variant reports into a
single report_mode="comparison" StatisticalReport with:
- variants — one VariantInfo per agent_model
- records — every variant's records, with
RecordResult.tagset to the full model identifier (e.g."gemini/gemini-2.5-flash", not the short label) so per-row provenance is unambiguous - metrics — every variant's metrics, renamed
<short-label>:<metric>and tagged with the full model invariant - gates — every variant's gates, metric_name renamed the same way
- comparisons — every non-baseline variant against the baseline, one per judge score, family-wise corrected
- power_analyses — one per binary-score comparison only;
post_hoc_power's MDE is a two-proportion quantity, so it shares units with the observed effect nowhere else - category_summaries — every variant's, metrics renamed and tagged
- failure_modes — every variant's,
failure_moderenamed the same way (FailureModeSummarycarries no variant field)
<short-label>:<metric> is one spelling shared by metrics and gates, and
that is what makes the merged report gateable: publish_report and
thresholds.yaml re-gating both index metrics by name and refuse to pick a
winner among duplicates, so bare per-variant metric names would leave the
report ungateable from either side.
baseline= names the reference arm; it defaults to the first model in
models (is_baseline=True).
Typical use:
from latent.flows.agent_garden_flow import agent_garden_flow
result = agent_garden_flow(
eval_data,
agent_factory_for_model=lambda m: lambda: make_agent(model=m),
judge=make_judge(),
models=["gemini/gemini-2.5-flash", "gemini/gemini-2.5-flash-lite"],
gates={"gemini-2.5-flash:accuracy": 0.6},
seed=42,
)
report = result["report"] # report_mode="comparison"
per_model = result["per_model"] # dict[str, agent_eval_flow result]
Functions¶
agent_garden_flow¶
agent_garden_flow(eval_data: pd.DataFrame, agent_factory_for_model: Callable[[str], Callable[[], Any]], judge: Any, models: list[str], baseline: str | None = None, question_column: str = 'question', expected_column: str | None = None, category_column: str | None = None, id_column: str | None = None, context_columns: list[str] | None = None, context_formatter: Callable[[dict[str, Any]], str] | None = None, gates: dict[str, float | GateSpec] | None = None, scorers: list[Callable[[pd.DataFrame, list[dict]], dict[str, MetricResult]]] | None = None, failure_taxonomy: dict[str, str] | None = None, failure_filter_fn: Callable | None = None, failure_timeout_s: float | None = None, failure_concurrency: int | None = None, concurrency: int = 1, judge_concurrency: int | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, adjust_method: str | None = 'holm', summary_model: str | None = None, summary_label: str | None = None) -> dict[str, Any]
Run agent_eval_flow per agent_model and merge into one comparison report.
Args:
eval_data: DataFrame with at least a question column.
agent_factory_for_model: Callable that takes a model name and returns
an agent factory (a zero-arg callable that produces a fresh
agent instance configured for that model). The double-callable
shape mirrors agent_eval_flow's agent_factory arg.
judge: A shared Judge — or list of Judges — applied to every variant's
outputs, exactly as agent_eval_flow takes it.
models: List of agent model identifiers to compare, each listed once.
Two models that shorten to the same label (openai/gpt-4o and
azure/gpt-4o) both keep their full identifier as the label.
baseline: The model every other variant is compared against. Defaults
to models[0]. Naming it explicitly is worth doing: the
baseline fixes the sign and the interpretation text of every
delta, so reordering models in a params file silently inverts
the whole report.
question_column / expected_column / category_column / id_column:
Forwarded to agent_eval_flow.
context_columns / context_formatter: Forwarded to agent_eval_flow.
gates: Quality thresholds. A <label>:<metric> key gates only that
variant — "gemini-2.5-flash:accuracy" is a floor on that model
alone. Passed programmatically, a bare key ("accuracy") gates
every variant; from a thresholds.yaml section a bare key is
rejected, because gates and metrics come out of the merge named
<label>:<metric> and the teardown re-gates the merged report by
name.
scorers: Forwarded to agent_eval_flow per variant — the same
deterministic scorers run against every model.
failure_taxonomy / failure_filter_fn / failure_timeout_s /
failure_concurrency / concurrency / judge_concurrency / n_resamples /
confidence_level / seed:
Forwarded to agent_eval_flow.
adjust_method: Family-wise correction over the comparison p-values —
one family spanning every (score, contender) pair, since they all
come off the same eval set. "holm" (default) controls FWER,
"fdr_bh" controls FDR, None leaves them uncorrected.
summary_model: Optional model identifier for client-side summary
generation. If set, generates a 3-5 sentence summary of the
combined report and attaches it to report.summary.
summary_label: Optional label passed to the summarizer. For a
garden run this should reflect that it's a comparison —
e.g. "Docs QA — gemini-flash vs gemini-flash-lite" —
so the summary text frames metrics in comparison terms
rather than as a single-agent baseline.
Returns: Dict with: - per_model: {model_name: agent_eval_flow result dict} - report: combined StatisticalReport with report_mode="comparison" - markdown: rendered markdown of the combined report - all_passed: True iff every variant's gates all passed
Raises: ValueError: models is empty or repeats a model, baseline is not among them, id_column repeats an id, or a gate key names an unknown variant.
Attributes¶
QUALIFIER¶
Separator between a variant label and a metric name.
A prefix, and this character, for two reasons measured rather than assumed:
MLflow validates metric names against ^[/\w.\- :]*$, so @ makes
log_metric raise and log_to_mlflow's catch-all swallows it — the whole
merged report writes no metrics, params or artifacts. And gates.slices
decodes a slice as {anchor}_{...}, a prefix relation, so a suffix
qualifier would stop every scores_by slice being exempt from threshold
coverage.