latent.stats.eval_comparison¶
High-level paired comparison over evaluation row dicts.
Eval pipelines typically produce one dict per row from each system being
compared, keyed by an id field. eval_comparison pairs those rows,
extracts the metric values via field name or callable, routes each metric
to the right paired test by score type via :func:latent.stats.report.analyze,
and applies family-wise correction across the metric set.
Use this whenever you have two systems scored on the same eval set and
want a publication-quality paired-comparison report. For ad-hoc / one-off
comparisons you can still call :func:compare_systems directly.
Functions¶
eval_comparison¶
eval_comparison(rows_a: list[dict[str, Any]], rows_b: list[dict[str, Any]], binary: MetricSpec | None = None, ordinal: MetricSpec | None = None, continuous: MetricSpec | None = None, lower_is_better: Iterable[str] | None = None, id_key: str = 'id', n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, adjust_method: str | None = 'holm') -> StatisticalReport
Paired comparison of two systems on the same eval set.
Pairs rows_a to rows_b by id_key, extracts each declared
metric via field name or callable, routes each metric to the right
paired test by score type (binary → McNemar, ordinal → Wilcoxon,
continuous → paired bootstrap; permutation fallback for n < 30), and
applies family-wise p-value correction across the metric set.
Each metric spec maps a result-name to either:
- a field name ("intent_answer_f1") — value pulled from
row["intent_answer_f1"].
- a callable row -> float|bool|None — for derived metrics like
correctness_pass = grade >= 3.
Returning None from a callable drops that row from that metric's
arrays only (other metrics are unaffected).
Args:
rows_a: Per-row dicts from system A (the baseline).
rows_b: Per-row dicts from system B (the candidate).
binary: Specs for binary metrics (e.g. correctness_pass).
ordinal: Specs for ordinal metrics (e.g. 1-4 letter grade).
continuous: Specs for continuous metrics (F1, latency, tokens).
lower_is_better: Set/iterable of metric names where a smaller
value is better (latency, tokens, cost). Only affects the
interpretation string in each ComparisonResult.
id_key: Field used to pair rows across systems. Defaults to
"id".
n_resamples: Bootstrap / permutation resamples.
confidence_level: Confidence level for the CIs.
seed: Random seed for reproducibility.
adjust_method: Family-wise correction across the comparison
p-values. "holm" (default) controls FWER; "fdr_bh"
controls FDR; None disables adjustment. Defaults to
"holm" here (unlike :func:analyze) because this entry
point is designed for the multi-metric case.
Returns: StatisticalReport with one comparison per declared metric. The single-system metrics on side A are also populated.
Raises: ValueError: If no rows share an id between the two sides, or if no metrics are declared. TypeError: If a metric spec is a bare string instead of a dict or iterable of names.