Skip to content

latent.stats.comparison

Two-system comparison and non-inferiority testing.

Reference: Delegates to statsmodels.stats.contingency_tables.mcnemar, scipy.stats.wilcoxon, scipy.stats.bootstrap, and latent.stats.bootstrap.paired_bootstrap_ci.

Functions

compare_systems

compare_systems(scores_a: np.ndarray, scores_b: np.ndarray, score_type: str = 'binary', n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, lower_is_better: bool = False) -> ComparisonResult

Paired comparison of System A vs System B on the same eval set.

Automatically selects the appropriate statistical test based on score_type and sample size.

Args: scores_a: Scores from system A. scores_b: Scores from system B (paired, same order). score_type: One of "binary", "ordinal", "continuous". n_resamples: Number of bootstrap / permutation resamples. confidence_level: Confidence level for the interval. seed: Random seed for reproducibility. lower_is_better: When True, a negative delta means B is better than A (e.g. latency, cost). Affects only the interpretation string in the returned result.

Returns: ComparisonResult with delta, CI, p-value, effect size, and interpretation.

Raises: ValueError: If arrays differ in length or score_type is unknown.

mcnemars_test

mcnemars_test(scores_a: np.ndarray, scores_b: np.ndarray) -> ComparisonResult

McNemar's test for paired binary data.

Reference: McNemar (1947), via statsmodels.stats.contingency_tables.mcnemar(exact=False, correction=True).

Tests whether the proportion of discordant pairs is symmetric, i.e. whether System A and System B make different types of errors at different rates.

Args: scores_a: Binary scores (0/1) from system A. scores_b: Binary scores (0/1) from system B (paired, same order).

Returns: ComparisonResult with the test statistic as delta and the chi-squared p-value.

Raises: ValueError: If arrays differ in length.

non_inferiority_test

non_inferiority_test(scores_a: np.ndarray, scores_b: np.ndarray, margin: float, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, lower_is_better: bool = False) -> GatingResult

Non-inferiority test: is System B not worse than A by more than margin?

Constructs a one-sided confidence interval for (B - A). The test passes if the lower bound of the CI exceeds -margin.

When lower_is_better=True the direction flips: the test passes if the upper bound of the CI is below margin (i.e. B is not worse than A by more than margin in the upward direction).

Reference: Bootstrap CI via scipy.stats.bootstrap(method='percentile').

Args: scores_a: Scores from system A (reference). scores_b: Scores from system B (candidate). margin: Maximum acceptable degradation (positive number). n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for the one-sided interval. seed: Random seed. lower_is_better: When True, flip the margin direction so that increases (not decreases) are penalised.

Returns: GatingResult with passed=True when B is non-inferior, stamped with the effective lower_is_better so renderers print the right direction.