Skip to content

latent.gates.thresholds

thresholds.yaml — the flow-scoped quality-gate lockfile.

config/thresholds.yaml holds one section per flow; a section's existence opts that flow into implicit gating (pre-built flows call :func:resolve_gates). Values are a bare float (gated at the default lower_ci strictness) or a {threshold, strictness} mapping parsed as :class:~latent.gates.gating.GateSpec.

Promotion is ratchet-only: :func:apply_promotion rewrites a threshold to GatingResult.promotable_to (computed in threshold_gate), never lower. Sections are keyed by the user's @flow name — pre-built flows invoked inside one flow share its section, so metric names within a flow are one namespace.

Functions

apply_promotion

apply_promotion(flow_name: str, metric: str, promotable_to: float) -> float | None

Ratchet one metric's threshold. Returns the new value, or None on no-op.

Never lowers: promotable_to <= current is a no-op (lowering a threshold is a deliberate manual edit, reviewed like any other change). Comment- preserving round-trip via ruamel; atomic replace.

bootstrap_section

bootstrap_section(flow_name: str, metrics: dict[str, float]) -> None

Append a new flow section (adoption / init). Never touches existing sections.

discover_report

discover_report(flow_name: str) -> tuple[Path, dict] | None

Newest gates-carrying report, or None. See :func:discover_reports.

discover_reports

discover_reports(flow_name: str, newer_than: float | None = None, run_id: str | None = None) -> list[tuple[Path, dict]]

Every JSON under the flow's output dir whose payload has a gates array.

Newest first by mtime — the single discovery mechanism shared by the CLI (status/promote/check/init) and the flow teardown. The search recurses: path: reports/eval.json is a documented catalog spelling, and a flat glob made such a report invisible to enforcement, which then failed the run for producing no report at all. Unrelated JSON is excluded by payload shape (a dict with a gates list), not by depth.

Two freshness rules that narrow together, because wall-clock is not an identity — but the absence of an identity is not a disqualification either:

  • run_id vetoes payloads stamped with a different id (:func:~latent.gates.publish.publish_report writes the stamp), and nothing else. A payload carrying this run's id is kept whatever its mtime. That is what makes two concurrent runs of one flow enforce their own reports, and what stops a backwards clock step from hiding a report the run just wrote.
  • newer_than keeps only files with mtime >= newer_than. It rules wherever identity is unavailable: no run_id argument (the CLI, direct callers), or a payload published where the run-id ContextVar does not reach (a process-based task runner, a subprocess worker). Unstamped is unknown provenance, not someone else's.

Parse failures are collected, not raised on sight. Beside at least one valid report a corrupt candidate is logged at WARNING and skipped, so a truncated leftover from a killed run cannot permanently fail every later run of the flow. With no valid report at all it raises a ValueError naming every corrupt candidate — "no report" would be a lie when the missing report may be the mangled one. Reports saved to absolute paths escape the output dir and are not discovered — status reports those flows as having no saved report.

load_thresholds

load_thresholds(flow_name: str | None = None) -> dict[str, GateSpec]

Load a flow's gates from thresholds.yaml, failing fast on missing config.

Explicit counterpart of :func:resolve_gates for custom flows that call analyze() directly. flow_name defaults to the active flow context.

promotable_gates

promotable_gates(payload: dict) -> list[dict]

The payload's gates that may ratchet the lockfile — [] for smoke evidence.

The one definition of "this run earned a promotion", shared by the teardown TTY prompt and latent thresholds promote. A sampled or gates-disabled run measured a subset (or gated nothing), so its promotable_to is not evidence a full run could defend — and the ratchet is one-way, so accepting it would permanently raise a threshold nothing ever measured.

prompt_thresholds

prompt_thresholds(flow_name: str, timeout: float = 30.0, newer_than: float | None = None) -> None

Interactive TTY prompt: adopt ungated headline metrics, then offer promotions.

Entry-point independent — called from the @flow teardown, so it fires identically under latent run, direct python3, or consumer wrappers. No-op unless stdin+stdout are a TTY, LATENT_PROMOTE_PROMPT != 0, and a report is discovered. newer_than scopes discovery itself (the teardown passes the flow start time), so a stale report is neither read nor parsed. Coverage and adoption span every fresh report — an orchestrator flow publishes several per run — while promotion reads the newest. Sampled or gates-disabled payloads print a skip note and offer nothing: smoke evidence must never bootstrap or ratchet thresholds. Adoption offers headlines only — slice metrics ({headline}_{...}) enter the lockfile solely by explicit hand edit.

read_sections

read_sections() -> dict[str, Any]

All raw flow sections from thresholds.yaml ({} when the file is absent).

regate

regate(payload: dict, section: dict[str, GateSpec]) -> list[GatingResult]

Re-gate a report payload's metrics against the current lockfile section.

The enforcement path's single source of verdicts. Gates baked into the payload are history — display and promotion evidence — and are ignored here: a report published when the threshold was 0.90 must fail once the section says 0.95, whatever the task cache did in between. Returns the live gate results in section order; a section key this payload has no metric for is simply absent. Staleness is not decidable here — in a multi-report flow a key is stale only when no payload reports it — so the teardown unions metric names across reports and judges it there.

Direction and strictness come from section alone — a payload carries no direction, and analyze's metric_directions channel does not reach this far. A ceiling metric must therefore be written {threshold: X, lower_is_better: true} in thresholds.yaml: as a bare float it re-gates as a floor, and a high leak rate would pass.

Raises: ValueError: The payload carries no metrics key — defaulting it would read as "zero gates ran" in the enforcement path — or a gated metric name appears more than once in it (per-variant reports name gates <variant>:<metric>; gate those names, not the bare metric). ValidationError: A payload metric is missing required fields.

resolve_gates

resolve_gates(gates: dict | None) -> dict[str, GateSpec] | dict | None

Implicit gate resolution for the pre-built flows.

An explicit gates argument wins verbatim (including {} to disable). Otherwise the active flow's thresholds.yaml section is used when one exists; a missing file/section means "not opted in" and resolves to None — but a malformed section still raises.

smoke_reason

smoke_reason(payload: dict) -> str | None

Why this payload must neither gate nor ratchet, or None if it may.

thresholds_path

thresholds_path() -> Path

uncovered_metrics

uncovered_metrics(flow_name: str, payloads: Sequence[dict]) -> list[str] | None

Metrics the flow's reports produced that thresholds.yaml doesn't gate.

Takes every discovered payload, not just the newest: an orchestrator flow publishes several reports per run, and a metric gated nowhere is uncovered whichever report named it. Names are unioned in first-seen order.

Returns None when the flow has no section at all (fully ungated), a list of metric names when the section exists but misses some, and [] when covered. Slice metrics are exempt: a name extending a covered name as {covered}_{...} (scores_by's naming) is gated only by an explicit section entry — covering the headline silences its whole slice family. With no section, slice-ness is judged against the reported names themselves, so an entirely ungated family still surfaces as ungated. An exempted name is invisible here by design; latent thresholds status renders it as a (slice — exempt) row so it is never silently dropped.