Skip to content

latent.stats.bootstrap

Bootstrap confidence intervals.

Reference: Delegates to scipy.stats.bootstrap for BCa and percentile methods. aligned_bootstrap_ci draws its own resamples instead — scipy's paired=True cannot stratify a draw, which ranking metrics on imbalanced data require.

Functions

aligned_bootstrap_ci

aligned_bootstrap_ci(arrays: np.ndarray = (), statistic: Callable[..., float], n_resamples: int = 10000, confidence_level: float = 0.95, stratify_by: np.ndarray | None = None, seed: int | None = None) -> MetricResult

Percentile bootstrap CI for a statistic of several row-aligned arrays.

Resamples index positions and indexes every array with the same draw, so row-wise pairing survives the resample. That is the shape ranking metrics need — AUROC and AUPRC are functions of (y_true, y_score) together, and bootstrap_ci cannot express them: it hands the statistic one resampled array, and closing over a fixed second array breaks as soon as the draw and the closed-over array have different lengths.

Pass stratify_by to resample within groups, holding each group's count fixed. For AUROC that is not a tuning knob but a correctness requirement: an unstratified draw from an imbalanced set can contain a single class, where the metric is undefined (roc_auc_score returns NaN and warns).

A resample whose statistic is non-finite is dropped and the percentiles are taken over the survivors — propagating the NaN would make the whole interval NaN, which says nothing, while a non-finite draw carries no information about the statistic either. Dropping conditions the interval on non-degeneracy, so it is bounded: below 90% surviving resamples this raises rather than return a differently-defined quantity under a 95% label, and any dropping at all is stamped into method as nowhere, because MetricResult.warnings was at the time rendered by no report surface at all. method reaches render_markdown and the MLflow params; render_text's metric table drops it too. Exceptions raised by statistic are not caught.

Args: arrays: One or more row-aligned arrays of equal length, passed to statistic positionally in this order. A single array is accepted and differs from bootstrap_ci only in method: this returns a percentile interval, bootstrap_ci a BCa one. statistic: Function of the resampled arrays returning a float, e.g. lambda y_true, y_score: roc_auc_score(y_true, y_score). n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. stratify_by: Optional per-row group labels. Each group is resampled to its own size, so every draw preserves the observed group counts. seed: Random seed for reproducibility.

Returns: MetricResult with point estimate and percentile CI bounds.

Raises: ValueError: If no arrays are given; the arrays (or stratify_by) differ in length; n_resamples is below 1; stratify_by leaves any row in no stratum (a NaN label equals nothing, so those rows would enter no draw while the point estimate still counts them); statistic is non-finite on the observed data; every resample produced a non-finite statistic; or fewer than 90% of them produced a finite one.

bootstrap_ci

bootstrap_ci(scores: np.ndarray, statistic: Callable[[np.ndarray], float] = np.mean, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> MetricResult

BCa bootstrap confidence interval for a statistic.

Reference: Efron (1987), via scipy.stats.bootstrap(method='BCa').

Args: scores: 1D array of scores. statistic: Function to compute the statistic (default: mean). n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility.

Returns: MetricResult with point estimate and CI bounds.

paired_bootstrap_ci

paired_bootstrap_ci(scores_a: np.ndarray, scores_b: np.ndarray, statistic: Callable[[np.ndarray], float] = np.mean, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> ComparisonResult

Paired bootstrap CI for difference between two systems.

Reference: via scipy.stats.bootstrap for CI, manual p-value calculation.

Args: scores_a: Scores from system A. scores_b: Scores from system B (same eval set, same order). statistic: Function to compute the statistic. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level. seed: Random seed.

Returns: ComparisonResult with delta, CI, and p-value.

quantile_bootstrap_ci

quantile_bootstrap_ci(scores: np.ndarray, quantile: float = 0.5, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> MetricResult

Bootstrap CI for an arbitrary quantile (e.g. p50, p95, p99).

Args: scores: 1D array of values. quantile: Quantile to estimate (0-1). E.g. 0.95 for p95. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility.

Returns: MetricResult with quantile estimate and BCa CI.