latent.stats.classification¶
Deterministic classification metrics with bootstrap CIs.
Reference: Delegates to scipy.stats.bootstrap for BCa intervals and
sklearn.metrics for classification metrics.
Functions¶
average_precision_ci¶
average_precision_ci(y_true: list | np.ndarray, y_score: list | np.ndarray, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> MetricResult
AUPRC (average precision) with a class-stratified paired bootstrap CI.
The estimand is average precision conditional on the observed class counts. Average precision is prevalence-dependent — its own baseline is the positive rate — so holding the counts fixed removes prevalence variance from the resample. When the labels are a random sample rather than fixed by design, the interval is therefore narrower than an unconditional one: measured over 300 trials at n=300 and prevalence 0.10, the stratified interval averaged ~9% narrower than the unstratified one. AUROC is prevalence-invariant and pays almost nothing for the same stratification, which is why it is not a knob there.
Note also that a percentile bootstrap for AP under-covers its nominal
level on imbalanced data (both variants measured ~89-92% at a nominal
95%), so a lower_ci gate on auprc is stricter than the label
implies.
Args: y_true: Binary ground-truth labels (0/1), both classes present. y_score: Predicted scores or probabilities for the positive class. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility.
Returns:
MetricResult named auprc, ready for threshold_gate like any
other metric. A minority class below ten members adds a warning: the
interval is driven by which of them the draw repeats.
Raises: ValueError: If the arrays differ in length, y_true is not binary 0/1 with both classes present, or either class has fewer than two members (a stratified resample of a single member is identical in every draw).
classification_metrics¶
classification_metrics(y_pred: list | np.ndarray, y_true: list | np.ndarray, labels: list[str] | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> dict[str, MetricResult]
Compute classification metrics with bootstrap CIs.
Binary classification: accuracy, precision, recall, F1, MCC. Multi-class: per-class precision/recall/F1, macro/micro/weighted F1.
Args: y_pred: Predicted labels. y_true: Actual labels. labels: Ordered list of label names. Inferred from data if None. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for CIs. seed: Random seed for reproducibility.
Returns: Dict mapping metric names to MetricResult objects.
confusion_matrix_with_ci¶
confusion_matrix_with_ci(y_pred: list | np.ndarray, y_true: list | np.ndarray, labels: list[str] | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> ConfusionMatrixResult
Confusion matrix with per-cell bootstrap CIs.
Args: y_pred: Predicted labels. y_true: Actual labels. labels: Ordered list of label names. Inferred from data if None. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for CIs. seed: Random seed for reproducibility.
Returns: ConfusionMatrixResult with matrix, ci_lower, ci_upper, labels.
intent_accuracy¶
intent_accuracy(expected_intents: list[str], actual_intents: list[str], confidence_level: float = 0.95) -> MetricResult
Classification accuracy for agent intent selection.
Measures whether the agent chose the correct intent (e.g. "escalate", "answer", "clarify") compared to the ground-truth expected intent.
Args: expected_intents: Ground-truth intents per interaction. actual_intents: Agent-selected intents per interaction. confidence_level: Confidence level for Wilson CI.
Returns: MetricResult with Wilson score interval.
intent_classification_report¶
intent_classification_report(expected: list[str], actual: list[str], prefix: str = '', n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> list[MetricResult]
Per-class precision/recall/F1 plus macro F1 for intent or category classification.
Returns a flat list of MetricResult objects with names like
{prefix}escalate_f1, {prefix}answer_precision, {prefix}f1_macro.
Args:
expected: Ground-truth labels per interaction.
actual: Predicted labels per interaction.
prefix: Optional prefix for metric names (e.g. "intent_").
n_resamples: Number of bootstrap resamples for CIs.
confidence_level: Confidence level for CIs.
seed: Random seed.
Returns: List of MetricResult objects.
outcome_accuracy¶
outcome_accuracy(expected_intents: list[str], actual_intents: list[str], answer_grades: list[str], pass_grades: set[str] | frozenset[str] = frozenset({'A', 'B'}), confidence_level: float = 0.95) -> MetricResult
Composite metric combining intent correctness with answer quality.
An interaction passes when:
- The agent selected the correct intent (e.g. expected escalation
and agent escalated), OR
- The agent produced a good answer (grade in pass_grades)
regardless of the expected intent.
An interaction fails when the agent selected the wrong intent AND
the answer quality is below pass_grades.
This avoids penalizing agents that answer well on questions where escalation was expected but not strictly necessary, and avoids penalizing correct escalations that receive low answer-quality grades.
Args: expected_intents: Ground-truth intents per interaction. actual_intents: Agent-selected intents per interaction. answer_grades: Quality grades per interaction (e.g. "A", "B", "C", "D"). pass_grades: Set of grades considered acceptable. confidence_level: Confidence level for Wilson CI.
Returns: MetricResult with Wilson score interval.
per_class_metrics¶
per_class_metrics(y_pred: list | np.ndarray, y_true: list | np.ndarray, labels: list[str] | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> dict[str, dict[str, MetricResult]]
Per-class precision, recall, and F1 with bootstrap CIs.
Args: y_pred: Predicted labels. y_true: Actual labels. labels: Ordered list of label names. Inferred from data if None. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for CIs. seed: Random seed for reproducibility.
Returns: Dict mapping class label -> {"precision": MetricResult, "recall": MetricResult, "f1": MetricResult}.
roc_auc_ci¶
roc_auc_ci(y_true: list | np.ndarray, y_score: list | np.ndarray, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> MetricResult
AUROC with a class-stratified paired bootstrap CI.
Args: y_true: Binary ground-truth labels (0/1), both classes present. y_score: Predicted scores or probabilities for the positive class. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility.
Returns:
MetricResult named auroc, ready for threshold_gate like any
other metric. A minority class below ten members adds a warning: the
interval is driven by which of them the draw repeats.
Raises: ValueError: If the arrays differ in length, y_true is not binary 0/1 with both classes present (AUROC is undefined otherwise), or either class has fewer than two members (a stratified resample of a single member is identical in every draw).