latent.stats.bootstrap¶
Bootstrap confidence intervals.
Reference: Delegates to scipy.stats.bootstrap for BCa and percentile methods.
aligned_bootstrap_ci draws its own resamples instead — scipy's paired=True
cannot stratify a draw, which ranking metrics on imbalanced data require.
Functions¶
aligned_bootstrap_ci¶
aligned_bootstrap_ci(arrays: np.ndarray = (), statistic: Callable[..., float], n_resamples: int = 10000, confidence_level: float = 0.95, stratify_by: np.ndarray | None = None, seed: int | None = None) -> MetricResult
Percentile bootstrap CI for a statistic of several row-aligned arrays.
Resamples index positions and indexes every array with the same draw, so
row-wise pairing survives the resample. That is the shape ranking metrics
need — AUROC and AUPRC are functions of (y_true, y_score) together, and
bootstrap_ci cannot express them: it hands the statistic one resampled
array, and closing over a fixed second array breaks as soon as the draw and
the closed-over array have different lengths.
Pass stratify_by to resample within groups, holding each group's count
fixed. For AUROC that is not a tuning knob but a correctness requirement:
an unstratified draw from an imbalanced set can contain a single class,
where the metric is undefined (roc_auc_score returns NaN and warns).
A resample whose statistic is non-finite is dropped and the percentiles are
taken over the survivors — propagating the NaN would make the whole interval
NaN, which says nothing, while a non-finite draw carries no information
about the statistic either. Dropping conditions the interval on
non-degeneracy, so it is bounded: below 90% surviving resamples this raises
rather than return a differently-defined quantity under a 95% label, and
any dropping at all is stamped into method as
nowhere, because MetricResult.warnings was at the time rendered
by no report surface at all. method reaches render_markdown and the
MLflow params; render_text's metric table drops it too.
Exceptions raised by statistic are not caught.
Args:
arrays: One or more row-aligned arrays of equal length, passed to
statistic positionally in this order. A single array is accepted
and differs from bootstrap_ci only in method: this returns a
percentile interval, bootstrap_ci a BCa one.
statistic: Function of the resampled arrays returning a float, e.g.
lambda y_true, y_score: roc_auc_score(y_true, y_score).
n_resamples: Number of bootstrap resamples.
confidence_level: Confidence level for the interval.
stratify_by: Optional per-row group labels. Each group is resampled to
its own size, so every draw preserves the observed group counts.
seed: Random seed for reproducibility.
Returns: MetricResult with point estimate and percentile CI bounds.
Raises: ValueError: If no arrays are given; the arrays (or stratify_by) differ in length; n_resamples is below 1; stratify_by leaves any row in no stratum (a NaN label equals nothing, so those rows would enter no draw while the point estimate still counts them); statistic is non-finite on the observed data; every resample produced a non-finite statistic; or fewer than 90% of them produced a finite one.
bootstrap_ci¶
bootstrap_ci(scores: np.ndarray, statistic: Callable[[np.ndarray], float] = np.mean, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> MetricResult
BCa bootstrap confidence interval for a statistic.
Reference: Efron (1987), via scipy.stats.bootstrap(method='BCa').
Args: scores: 1D array of scores. statistic: Function to compute the statistic (default: mean). n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility.
Returns: MetricResult with point estimate and CI bounds.
paired_bootstrap_ci¶
paired_bootstrap_ci(scores_a: np.ndarray, scores_b: np.ndarray, statistic: Callable[[np.ndarray], float] = np.mean, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> ComparisonResult
Paired bootstrap CI for difference between two systems.
Reference: via scipy.stats.bootstrap for CI, manual p-value calculation.
Args: scores_a: Scores from system A. scores_b: Scores from system B (same eval set, same order). statistic: Function to compute the statistic. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level. seed: Random seed.
Returns: ComparisonResult with delta, CI, and p-value.
quantile_bootstrap_ci¶
quantile_bootstrap_ci(scores: np.ndarray, quantile: float = 0.5, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None) -> MetricResult
Bootstrap CI for an arbitrary quantile (e.g. p50, p95, p99).
Args: scores: 1D array of values. quantile: Quantile to estimate (0-1). E.g. 0.95 for p95. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility.
Returns: MetricResult with quantile estimate and BCa CI.