Skip to content

latent.flows.agent_eval_flow

agent_eval_flow — Full agent evaluation: inference → judge → scorers → report.

Functions

agent_eval_flow

agent_eval_flow(eval_data: pd.DataFrame, agent_factory: Callable[[], Any] | None = None, judge: Any = None, agent_factory_for_row: Callable[[dict[str, Any]], Any] | None = None, question_column: str = 'question', expected_column: str | None = None, category_column: str | None = None, id_column: str | None = None, context_columns: list[str] | None = None, context_formatter: Callable[[dict[str, Any]], str] | None = None, gates: dict[str, float | GateSpec] | None = None, scorers: list[Callable[[pd.DataFrame, list[dict]], dict[str, MetricResult]]] | None = None, post_inference: Callable[[pd.DataFrame], pd.DataFrame | Awaitable[pd.DataFrame]] | None = None, failure_taxonomy: dict[str, str] | None = None, failure_filter_fn: Callable | None = None, failure_timeout_s: float | None = None, failure_concurrency: int | None = None, concurrency: int = 1, judge_concurrency: int | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, summary_model: str | None = None, summary_label: str | None = None) -> dict[str, Any]

Full agent evaluation pipeline: inference → judge → scorers → report.

Orchestrates: 1. Run agent on each row via agent_inference_flow. 2. Merge inference outputs into eval_data. 3. Score with LLM judge via judge_flow. 4. Run optional deterministic scorers. 5. Build records, category breakdowns, and failure mode classification.

Args: eval_data: DataFrame with at least a question column. agent_factory: Callable returning a BaseAgent instance. Use this when every row gets the same kind of agent. agent_factory_for_row: Callable(row_dict) -> BaseAgent, for when the agent depends on the row — a per-channel context, a persona, a conversation history. Mutually exclusive with agent_factory. Without it, a caller whose agent varies per row has to group the frame and run one inference pass per group. judge: A Judge with .output_type, or a list of them. Each judge scores every eligible row and its output columns are concatenated onto one scored frame; two judges sharing a column name are rejected unless one carries a column_prefix. post_inference: Optional hook applied to the merged frame after inference and before judging. This is the seam for a column a judge needs that only exists once the agent has run — a faithfulness judge's retrieved_context, say. May be async, so it can await a retriever. Must return the frame it was given with columns added — df.assign(...) — never a projection, a filter or a re-sort of it: every scorer pairs scored_data with the inference_results list by position, and the later steps read output plus whichever question/expected/category/id columns were declared. Both are enforced. question_column: Column with questions/prompts. expected_column: Column with ground truth answers. category_column: Column with category labels. id_column: Column with row IDs. context_columns: Columns to include as context for the agent. context_formatter: Custom formatter for context columns. gates: Quality thresholds {"metric_name": threshold-or-GateSpec}. scorers: Optional deterministic scoring functions. Each receives (scored_data DataFrame, inference_results list) and returns a dict of MetricResult. failure_taxonomy: {"mode": "description"} for failure classification. failure_filter_fn: Predicate selecting records for failure classification. concurrency: Parallel agent instances. judge_concurrency: Rows scored concurrently by the judge. None (default) leaves judge_flow on its own default. Separate from concurrency because an agent instance is usually far more expensive per row than a judge call. failure_concurrency: Records classified concurrently in the failure step. None (default) leaves classify_failure_modes serial. It is deliberately not tied to concurrency: the classifier runs against its own model, so an agent-instance budget says nothing about the fan-out that model's provider will take. failure_timeout_s: Deadline for one record's failure classification; on expiry that record's mode stays unset and the rest continue. n_resamples: Bootstrap resamples. confidence_level: CI confidence level. seed: Random seed.

Returns: Dict with: - inference_results: list[dict] from agent inference. - scored_data: pd.DataFrame with judge scores. - report: StatisticalReport with records, categories, failure_modes. - markdown: str rendered report. - all_passed: bool for quality gates. - score_types: judge_flow's merged {column: score type}, every judge's column_prefix applied. Re-exposed so a caller that needs the scored column names — agent_garden_flow pairing them across variants — reads them instead of deriving them from one judge.