latent.flows.agent_eval_flow¶
agent_eval_flow — Full agent evaluation: inference → judge → scorers → report.
Functions¶
agent_eval_flow¶
agent_eval_flow(eval_data: pd.DataFrame, agent_factory: Callable[[], Any] | None = None, judge: Any = None, agent_factory_for_row: Callable[[dict[str, Any]], Any] | None = None, question_column: str = 'question', expected_column: str | None = None, category_column: str | None = None, id_column: str | None = None, context_columns: list[str] | None = None, context_formatter: Callable[[dict[str, Any]], str] | None = None, gates: dict[str, float | GateSpec] | None = None, scorers: list[Callable[[pd.DataFrame, list[dict]], dict[str, MetricResult]]] | None = None, post_inference: Callable[[pd.DataFrame], pd.DataFrame | Awaitable[pd.DataFrame]] | None = None, failure_taxonomy: dict[str, str] | None = None, failure_filter_fn: Callable | None = None, failure_timeout_s: float | None = None, failure_concurrency: int | None = None, concurrency: int = 1, judge_concurrency: int | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, summary_model: str | None = None, summary_label: str | None = None) -> dict[str, Any]
Full agent evaluation pipeline: inference → judge → scorers → report.
Orchestrates: 1. Run agent on each row via agent_inference_flow. 2. Merge inference outputs into eval_data. 3. Score with LLM judge via judge_flow. 4. Run optional deterministic scorers. 5. Build records, category breakdowns, and failure mode classification.
Args:
eval_data: DataFrame with at least a question column.
agent_factory: Callable returning a BaseAgent instance. Use this when
every row gets the same kind of agent.
agent_factory_for_row: Callable(row_dict) -> BaseAgent, for when the
agent depends on the row — a per-channel context, a persona, a
conversation history. Mutually exclusive with agent_factory.
Without it, a caller whose agent varies per row has to group the
frame and run one inference pass per group.
judge: A Judge with .output_type, or a list of them. Each judge
scores every eligible row and its output columns are concatenated
onto one scored frame; two judges sharing a column name are
rejected unless one carries a column_prefix.
post_inference: Optional hook applied to the merged frame after
inference and before judging. This is the seam for a column a
judge needs that only exists once the agent has run — a
faithfulness judge's retrieved_context, say. May be async, so
it can await a retriever. Must return the frame it was given with
columns added — df.assign(...) — never a projection, a filter
or a re-sort of it: every scorer pairs scored_data with the
inference_results list by position, and the later steps read
output plus whichever question/expected/category/id columns
were declared. Both are enforced.
question_column: Column with questions/prompts.
expected_column: Column with ground truth answers.
category_column: Column with category labels.
id_column: Column with row IDs.
context_columns: Columns to include as context for the agent.
context_formatter: Custom formatter for context columns.
gates: Quality thresholds {"metric_name": threshold-or-GateSpec}.
scorers: Optional deterministic scoring functions. Each receives
(scored_data DataFrame, inference_results list) and returns
a dict of MetricResult.
failure_taxonomy: {"mode": "description"} for failure classification.
failure_filter_fn: Predicate selecting records for failure classification.
concurrency: Parallel agent instances.
judge_concurrency: Rows scored concurrently by the judge. None
(default) leaves judge_flow on its own default. Separate from
concurrency because an agent instance is usually far more
expensive per row than a judge call.
failure_concurrency: Records classified concurrently in the failure
step. None (default) leaves classify_failure_modes serial.
It is deliberately not tied to concurrency: the classifier runs
against its own model, so an agent-instance budget says nothing
about the fan-out that model's provider will take.
failure_timeout_s: Deadline for one record's failure classification;
on expiry that record's mode stays unset and the rest continue.
n_resamples: Bootstrap resamples.
confidence_level: CI confidence level.
seed: Random seed.
Returns:
Dict with:
- inference_results: list[dict] from agent inference.
- scored_data: pd.DataFrame with judge scores.
- report: StatisticalReport with records, categories, failure_modes.
- markdown: str rendered report.
- all_passed: bool for quality gates.
- score_types: judge_flow's merged {column: score type}, every judge's
column_prefix applied. Re-exposed so a caller that needs the
scored column names — agent_garden_flow pairing them across
variants — reads them instead of deriving them from one judge.