Skip to content

latent.flows.judge_flow

judge_flow — Score a dataset with one or more Judges, analyze, and gate.

Each Judge supplies a Pydantic output_type; rows are scored concurrently via asyncio.gather (bounded by an asyncio.Semaphore) and the judge's output fields are appended to a copy of the input DataFrame. Multiple judges run sequentially across judges but concurrent across rows. Eligibility and row-mapping per judge let one Judge be applied to a subset of rows (e.g. FaithfulnessJudge skips rows without context).

Functions

judge_flow

judge_flow(eval_data: pd.DataFrame, judges: Any, concurrency: int = 10, gates: dict[str, float | GateSpec] | None = None, score_types: dict[str, str] | None = None, rubrics: dict | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, summary_model: str | None = None, summary_label: str | None = None) -> dict[str, Any]

Score a dataset with one or more Judges, run statistical analysis, check gates.

Args: eval_data: DataFrame with columns matching each judge's prompt_template (or with the columns each judge's row_mapper translates from). judges: A Judge[T] instance OR a list of Judge instances. When a list, each judge scores rows in parallel (per-judge) and their output columns are concatenated onto a single scored_df. Judges with an is_eligible(row) -> bool method are only applied to rows where that returns True. concurrency: Number of rows scored concurrently per judge (bound of an asyncio.Semaphore). gates: Optional dict mapping metric name to threshold (float or GateSpec). score_types: Optional override for score type detection. rubrics: Optional override for rubric detection. n_resamples: Bootstrap resamples for analysis. confidence_level: CI confidence level. seed: Random seed for reproducibility.

Returns: Dict with keys: scored_data, report, markdown, all_passed, score_types (column -> score type, prefixes applied).