latent.stats.comparison¶
Two-system comparison and non-inferiority testing.
Reference: Delegates to statsmodels.stats.contingency_tables.mcnemar,
scipy.stats.wilcoxon, scipy.stats.bootstrap, and
latent.stats.bootstrap.paired_bootstrap_ci.
Functions¶
compare_systems¶
compare_systems(scores_a: np.ndarray, scores_b: np.ndarray, score_type: str = 'binary', n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, lower_is_better: bool = False) -> ComparisonResult
Paired comparison of System A vs System B on the same eval set.
Automatically selects the appropriate statistical test based on
score_type and sample size.
Args:
scores_a: Scores from system A.
scores_b: Scores from system B (paired, same order).
score_type: One of "binary", "ordinal", "continuous".
n_resamples: Number of bootstrap / permutation resamples.
confidence_level: Confidence level for the interval.
seed: Random seed for reproducibility.
lower_is_better: When True, a negative delta means B is
better than A (e.g. latency, cost). Affects only the
interpretation string in the returned result.
Returns: ComparisonResult with delta, CI, p-value, effect size, and interpretation.
Raises: ValueError: If arrays differ in length or score_type is unknown.
mcnemars_test¶
McNemar's test for paired binary data.
Reference: McNemar (1947), via
statsmodels.stats.contingency_tables.mcnemar(exact=False, correction=True).
Tests whether the proportion of discordant pairs is symmetric, i.e. whether System A and System B make different types of errors at different rates.
Args: scores_a: Binary scores (0/1) from system A. scores_b: Binary scores (0/1) from system B (paired, same order).
Returns: ComparisonResult with the test statistic as delta and the chi-squared p-value.
Raises: ValueError: If arrays differ in length.
non_inferiority_test¶
non_inferiority_test(scores_a: np.ndarray, scores_b: np.ndarray, margin: float, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, lower_is_better: bool = False) -> GatingResult
Non-inferiority test: is System B not worse than A by more than margin?
Constructs a one-sided confidence interval for (B - A). The test
passes if the lower bound of the CI exceeds -margin.
When lower_is_better=True the direction flips: the test passes
if the upper bound of the CI is below margin (i.e. B is not
worse than A by more than margin in the upward direction).
Reference: Bootstrap CI via scipy.stats.bootstrap(method='percentile').
Args:
scores_a: Scores from system A (reference).
scores_b: Scores from system B (candidate).
margin: Maximum acceptable degradation (positive number).
n_resamples: Number of bootstrap resamples.
confidence_level: Confidence level for the one-sided interval.
seed: Random seed.
lower_is_better: When True, flip the margin direction so that
increases (not decreases) are penalised.
Returns:
GatingResult with passed=True when B is non-inferior, stamped with
the effective lower_is_better so renderers print the right
direction.