latent.gates.slices¶
Grouped score arrays for sliced metric reporting.
scores_by replaces the per-flow groupby ritual: one scored DataFrame
in, the {metric_name: ndarray} mapping that analyze consumes out —
the overall (headline) array plus one array per group value per by
column. The overall array is included so the headline metric name and its
slice names come from one place and cannot drift in spelling.
This module defines the {prefix}_{col}_{val} slice-name convention, so
both halves of it live here: :func:scores_by encodes it, :func:is_slice
decodes it lexically. Every consumer decodes — slice-ness is asked about names
read back off a published payload or a thresholds.yaml section, never about
the mapping that minted them.
Functions¶
is_slice¶
True when name extends some anchor as {anchor}_{...} (never itself).
The decoder for :func:scores_by's naming, for consumers that hold only
names — a published payload, a thresholds.yaml section. Being lexical it
cannot tell a real slice from a metric that merely reads like one
(f1_macro beside f1, latency_p95 beside latency), so do not
name a headline metric after another one.
scores_by¶
scores_by(df: pd.DataFrame, value: str, by: str | Sequence[str], prefix: str | None = None) -> dict[str, np.ndarray]
Build overall plus per-slice score arrays from a scored DataFrame.
Args:
df: Scored rows, one per case.
value: Column holding the per-row score. Bool columns coerce to
0/1 floats.
by: Column name(s) to slice by. Each column slices independently —
no cross-product. A bare string means that one column. NaN
group values become a visible "nan" slice (rows are kept,
not dropped).
prefix: Name for the overall array and stem for slice names.
Defaults to value.
Returns:
Insertion-ordered mapping: the overall float array under prefix
first, then one entry per group value per by column (columns in
call order) named f"{prefix}_{col}_{val}", with group values
rendered via astype(str) and sorted lexicographically within each
column.
Raises:
ValueError: If the value column contains NaN (filter unscored
rows first — the partition stays visible in consumer code), or
if two by columns mint the same slice name.