Skip to content

latent.gates.slices

Grouped score arrays for sliced metric reporting.

scores_by replaces the per-flow groupby ritual: one scored DataFrame in, the {metric_name: ndarray} mapping that analyze consumes out — the overall (headline) array plus one array per group value per by column. The overall array is included so the headline metric name and its slice names come from one place and cannot drift in spelling.

This module defines the {prefix}_{col}_{val} slice-name convention, so both halves of it live here: :func:scores_by encodes it, :func:is_slice decodes it lexically. Every consumer decodes — slice-ness is asked about names read back off a published payload or a thresholds.yaml section, never about the mapping that minted them.

Functions

is_slice

is_slice(name: str, anchors: Collection[str]) -> bool

True when name extends some anchor as {anchor}_{...} (never itself).

The decoder for :func:scores_by's naming, for consumers that hold only names — a published payload, a thresholds.yaml section. Being lexical it cannot tell a real slice from a metric that merely reads like one (f1_macro beside f1, latency_p95 beside latency), so do not name a headline metric after another one.

scores_by

scores_by(df: pd.DataFrame, value: str, by: str | Sequence[str], prefix: str | None = None) -> dict[str, np.ndarray]

Build overall plus per-slice score arrays from a scored DataFrame.

Args: df: Scored rows, one per case. value: Column holding the per-row score. Bool columns coerce to 0/1 floats. by: Column name(s) to slice by. Each column slices independently — no cross-product. A bare string means that one column. NaN group values become a visible "nan" slice (rows are kept, not dropped). prefix: Name for the overall array and stem for slice names. Defaults to value.

Returns: Insertion-ordered mapping: the overall float array under prefix first, then one entry per group value per by column (columns in call order) named f"{prefix}_{col}_{val}", with group values rendered via astype(str) and sorted lexicographically within each column.

Raises: ValueError: If the value column contains NaN (filter unscored rows first — the partition stays visible in consumer code), or if two by columns mint the same slice name.