Skip to content

latent.stats.eval_comparison

High-level paired comparison over evaluation row dicts.

Eval pipelines typically produce one dict per row from each system being compared, keyed by an id field. eval_comparison pairs those rows, extracts the metric values via field name or callable, routes each metric to the right paired test by score type via :func:latent.stats.report.analyze, and applies family-wise correction across the metric set.

Use this whenever you have two systems scored on the same eval set and want a publication-quality paired-comparison report. For ad-hoc / one-off comparisons you can still call :func:compare_systems directly.

Functions

eval_comparison

eval_comparison(rows_a: list[dict[str, Any]], rows_b: list[dict[str, Any]], binary: MetricSpec | None = None, ordinal: MetricSpec | None = None, continuous: MetricSpec | None = None, lower_is_better: Iterable[str] | None = None, id_key: str = 'id', n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, adjust_method: str | None = 'holm') -> StatisticalReport

Paired comparison of two systems on the same eval set.

Pairs rows_a to rows_b by id_key, extracts each declared metric via field name or callable, routes each metric to the right paired test by score type (binary → McNemar, ordinal → Wilcoxon, continuous → paired bootstrap; permutation fallback for n < 30), and applies family-wise p-value correction across the metric set.

Each metric spec maps a result-name to either: - a field name ("intent_answer_f1") — value pulled from row["intent_answer_f1"]. - a callable row -> float|bool|None — for derived metrics like correctness_pass = grade >= 3.

Returning None from a callable drops that row from that metric's arrays only (other metrics are unaffected).

Args: rows_a: Per-row dicts from system A (the baseline). rows_b: Per-row dicts from system B (the candidate). binary: Specs for binary metrics (e.g. correctness_pass). ordinal: Specs for ordinal metrics (e.g. 1-4 letter grade). continuous: Specs for continuous metrics (F1, latency, tokens). lower_is_better: Set/iterable of metric names where a smaller value is better (latency, tokens, cost). Only affects the interpretation string in each ComparisonResult. id_key: Field used to pair rows across systems. Defaults to "id". n_resamples: Bootstrap / permutation resamples. confidence_level: Confidence level for the CIs. seed: Random seed for reproducibility. adjust_method: Family-wise correction across the comparison p-values. "holm" (default) controls FWER; "fdr_bh" controls FDR; None disables adjustment. Defaults to "holm" here (unlike :func:analyze) because this entry point is designed for the multi-metric case.

Returns: StatisticalReport with one comparison per declared metric. The single-system metrics on side A are also populated.

Raises: ValueError: If no rows share an id between the two sides, or if no metrics are declared. TypeError: If a metric spec is a bare string instead of a dict or iterable of names.

Attributes

Extractor

MetricSpec