Skip to content

latent.scores.schema_leak

Schema-leak scorer for agent_eval_flow.

Wraps the shared detector (guardrails.schema_leak) in the (df, records) -> dict[str, MetricResult] shape the eval flow expects. Vocab-bound closure — the caller builds a SchemaVocab from their backend and hands it in; latent-py never opens a DB.

Metrics emitted:

  • schema_leak_clean — deterministic 0/1 with degenerate CI. 1.0 iff zero strict leaks across all scorable rows AND at least one row was scorable (fail-closed on all-unscorable). Wire into agent_eval_flow(..., gates= {"schema_leak_clean": 0.5}) — the 0.5 threshold with strict-> gating passes clean=1.0 and fails leak=0.0.
  • schema_leak_count — number of leaking rows (diagnostic, not gated).
  • schema_leak_rate — leaking / scorable proportion (diagnostic).
  • schema_leak_scorable_rate — scorable / total rows, surfacing low coverage runs.

MetricResult.warnings on schema_leak_clean carries the leak evidence in-report (bounded — first N leaking rows, prefixed with the true total, per- item truncated) so all_passed=False is diagnosable without re-running or digging into the df. This is the primary evidence surface because arbitrary columns added to the scored df are not propagated to RecordResult by build_records_from_df (only judge-declared score columns are).

Functions

make_schema_leak_scorer

make_schema_leak_scorer(vocab: SchemaVocab, output_column: str = 'output', is_scorable: Callable[[Any], bool] = _is_nonempty_string, id_column: str = 'id') -> Callable[[pd.DataFrame, list[dict]], dict[str, MetricResult]]

Return an agent_eval_flow scorer that gates on schema leaks.

Args: vocab: The DB identifiers to detect leaks against. Built project-side from the backend's schema; latent-py never opens a DB. output_column: Column on the scored df carrying the agent's answer. Default "output" matches latent's canonical field. is_scorable: Predicate deciding whether a row's value is worth scoring. Default is generic non-empty-string. Projects with an error-marker convention (e.g. yadata's "ERROR:" sentinel) can pass a stricter predicate. id_column: Column carrying row identifiers, used in the leak warnings. Falls back to positional index if absent.

Raises: KeyError: at scoring time if output_column is missing from the df. This is a wiring bug (categorically different from a leak being found) and must not masquerade as a failed metric — the security gate is not the place to surface typos in flow config.