latent.flows.judge_flow¶
judge_flow — Score a dataset with one or more Judges, analyze, and gate.
Each Judge supplies a Pydantic output_type; rows are scored concurrently
via asyncio.gather (bounded by an asyncio.Semaphore) and the judge's
output fields are appended to a copy of the input DataFrame. Multiple judges
run sequentially across judges but concurrent across rows. Eligibility and
row-mapping per judge let one Judge be applied to a subset of rows (e.g.
FaithfulnessJudge skips rows without context).
Functions¶
judge_flow¶
judge_flow(eval_data: pd.DataFrame, judges: Any, concurrency: int = 10, gates: dict[str, float | GateSpec] | None = None, score_types: dict[str, str] | None = None, rubrics: dict | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, summary_model: str | None = None, summary_label: str | None = None) -> dict[str, Any]
Score a dataset with one or more Judges, run statistical analysis, check gates.
Args:
eval_data: DataFrame with columns matching each judge's prompt_template
(or with the columns each judge's row_mapper translates from).
judges: A Judge[T] instance OR a list of Judge instances. When a list,
each judge scores rows in parallel (per-judge) and their output
columns are concatenated onto a single scored_df. Judges with
an is_eligible(row) -> bool method are only applied to rows
where that returns True.
concurrency: Number of rows scored concurrently per judge (bound of
an asyncio.Semaphore).
gates: Optional dict mapping metric name to threshold (float or GateSpec).
score_types: Optional override for score type detection.
rubrics: Optional override for rubric detection.
n_resamples: Bootstrap resamples for analysis.
confidence_level: CI confidence level.
seed: Random seed for reproducibility.
Returns: Dict with keys: scored_data, report, markdown, all_passed, score_types (column -> score type, prefixes applied).