Skip to content

latent.stats.classification

Deterministic classification metrics with bootstrap CIs.

Reference: Delegates to scipy.stats.bootstrap for BCa intervals and sklearn.metrics for classification metrics.

Functions

average_precision_ci

average_precision_ci(y_true: list | np.ndarray, y_score: list | np.ndarray, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> MetricResult

AUPRC (average precision) with a class-stratified paired bootstrap CI.

The estimand is average precision conditional on the observed class counts. Average precision is prevalence-dependent — its own baseline is the positive rate — so holding the counts fixed removes prevalence variance from the resample. When the labels are a random sample rather than fixed by design, the interval is therefore narrower than an unconditional one: measured over 300 trials at n=300 and prevalence 0.10, the stratified interval averaged ~9% narrower than the unstratified one. AUROC is prevalence-invariant and pays almost nothing for the same stratification, which is why it is not a knob there.

Note also that a percentile bootstrap for AP under-covers its nominal level on imbalanced data (both variants measured ~89-92% at a nominal 95%), so a lower_ci gate on auprc is stricter than the label implies.

Args: y_true: Binary ground-truth labels (0/1), both classes present. y_score: Predicted scores or probabilities for the positive class. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility.

Returns: MetricResult named auprc, ready for threshold_gate like any other metric. A minority class below ten members adds a warning: the interval is driven by which of them the draw repeats.

Raises: ValueError: If the arrays differ in length, y_true is not binary 0/1 with both classes present, or either class has fewer than two members (a stratified resample of a single member is identical in every draw).

classification_metrics

classification_metrics(y_pred: list | np.ndarray, y_true: list | np.ndarray, labels: list[str] | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> dict[str, MetricResult]

Compute classification metrics with bootstrap CIs.

Binary classification: accuracy, precision, recall, F1, MCC. Multi-class: per-class precision/recall/F1, macro/micro/weighted F1.

Args: y_pred: Predicted labels. y_true: Actual labels. labels: Ordered list of label names. Inferred from data if None. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for CIs. seed: Random seed for reproducibility.

Returns: Dict mapping metric names to MetricResult objects.

confusion_matrix_with_ci

confusion_matrix_with_ci(y_pred: list | np.ndarray, y_true: list | np.ndarray, labels: list[str] | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> ConfusionMatrixResult

Confusion matrix with per-cell bootstrap CIs.

Args: y_pred: Predicted labels. y_true: Actual labels. labels: Ordered list of label names. Inferred from data if None. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for CIs. seed: Random seed for reproducibility.

Returns: ConfusionMatrixResult with matrix, ci_lower, ci_upper, labels.

intent_accuracy

intent_accuracy(expected_intents: list[str], actual_intents: list[str], confidence_level: float = 0.95) -> MetricResult

Classification accuracy for agent intent selection.

Measures whether the agent chose the correct intent (e.g. "escalate", "answer", "clarify") compared to the ground-truth expected intent.

Args: expected_intents: Ground-truth intents per interaction. actual_intents: Agent-selected intents per interaction. confidence_level: Confidence level for Wilson CI.

Returns: MetricResult with Wilson score interval.

intent_classification_report

intent_classification_report(expected: list[str], actual: list[str], prefix: str = '', n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> list[MetricResult]

Per-class precision/recall/F1 plus macro F1 for intent or category classification.

Returns a flat list of MetricResult objects with names like {prefix}escalate_f1, {prefix}answer_precision, {prefix}f1_macro.

Args: expected: Ground-truth labels per interaction. actual: Predicted labels per interaction. prefix: Optional prefix for metric names (e.g. "intent_"). n_resamples: Number of bootstrap resamples for CIs. confidence_level: Confidence level for CIs. seed: Random seed.

Returns: List of MetricResult objects.

outcome_accuracy

outcome_accuracy(expected_intents: list[str], actual_intents: list[str], answer_grades: list[str], pass_grades: set[str] | frozenset[str] = frozenset({'A', 'B'}), confidence_level: float = 0.95) -> MetricResult

Composite metric combining intent correctness with answer quality.

An interaction passes when: - The agent selected the correct intent (e.g. expected escalation and agent escalated), OR - The agent produced a good answer (grade in pass_grades) regardless of the expected intent.

An interaction fails when the agent selected the wrong intent AND the answer quality is below pass_grades.

This avoids penalizing agents that answer well on questions where escalation was expected but not strictly necessary, and avoids penalizing correct escalations that receive low answer-quality grades.

Args: expected_intents: Ground-truth intents per interaction. actual_intents: Agent-selected intents per interaction. answer_grades: Quality grades per interaction (e.g. "A", "B", "C", "D"). pass_grades: Set of grades considered acceptable. confidence_level: Confidence level for Wilson CI.

Returns: MetricResult with Wilson score interval.

per_class_metrics

per_class_metrics(y_pred: list | np.ndarray, y_true: list | np.ndarray, labels: list[str] | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> dict[str, dict[str, MetricResult]]

Per-class precision, recall, and F1 with bootstrap CIs.

Args: y_pred: Predicted labels. y_true: Actual labels. labels: Ordered list of label names. Inferred from data if None. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for CIs. seed: Random seed for reproducibility.

Returns: Dict mapping class label -> {"precision": MetricResult, "recall": MetricResult, "f1": MetricResult}.

roc_auc_ci

roc_auc_ci(y_true: list | np.ndarray, y_score: list | np.ndarray, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> MetricResult

AUROC with a class-stratified paired bootstrap CI.

Args: y_true: Binary ground-truth labels (0/1), both classes present. y_score: Predicted scores or probabilities for the positive class. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility.

Returns: MetricResult named auroc, ready for threshold_gate like any other metric. A minority class below ten members adds a warning: the interval is driven by which of them the draw repeats.

Raises: ValueError: If the arrays differ in length, y_true is not binary 0/1 with both classes present (AUROC is undefined otherwise), or either class has fewer than two members (a stratified resample of a single member is identical in every draw).