Skip to content

latent.flows.eval_report_flow

eval_report_flow — Aggregate evaluation results into a StatisticalReport.

Functions

eval_report_flow

eval_report_flow(eval_results: list[dict], score_columns: dict[str, str] | None = None, question_column: str = 'question', output_column: str = 'output', expected_column: str | None = None, category_column: str | None = None, id_column: str | None = None, failure_taxonomy: dict[str, str] | None = None, gates: dict[str, float | GateSpec] | None = None, baseline_results: list[dict] | None = None, comparison_results: dict[str, list[dict]] | None = None, adjust_method: str | None = 'holm', n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> dict[str, Any]

Build a complete StatisticalReport from evaluation results.

Aggregates metrics with CIs, category breakdowns, failure modes, quality gates, optional drift detection, and optional system comparisons.

Args: eval_results: List of dicts with evaluation results (one per row). score_columns: {"column_name": "binary"|"ordinal"|"continuous"} mapping of which columns contain scores and their types. If None, auto-detects numeric columns. question_column: Column with input questions. output_column: Column with agent outputs. expected_column: Column with expected/ground truth answers. category_column: Column with category labels. id_column: Column with record IDs. failure_taxonomy: {"mode": "description"} dict for failure mode summaries. gates: Quality thresholds {"metric_name": threshold}. baseline_results: Previous eval results for drift detection. comparison_results: {"system_name": [results]} for system comparisons. adjust_method: Family-wise correction across every comparison this report publishes ("holm", "fdr_bh", ...). None publishes raw per-comparison p-values, which over-report significance once the family exceeds one. n_resamples: Bootstrap resamples. confidence_level: CI confidence level. seed: Random seed.

Returns: Dict with: - report: StatisticalReport - markdown: str - all_passed: bool