latent.stats.report¶
Report orchestrator and method explanation (FR-8.1 + FR-8.4).
Functions¶
adjust_comparisons¶
adjust_comparisons(comparisons: list[ComparisonResult], method: AdjustMethod) -> list[ComparisonResult]
Apply a family-wise correction to a set of comparison p-values.
The one place the corrected p-value, the (adj:<method>) method suffix
and the rewritten interpretation tail are produced together. Every builder
(compare_systems and friends) embeds the verdict and raw p-value at the
tail of interpretation (e.g. "(statistically significant, p=0.0300)");
after correction the p-value changes, sometimes crossing the significance
threshold, so the text is patched to match. Without it callers read
p_value=0.09 beside "...significant, p=0.0300".
A single comparison is returned untouched — there is no family to correct. method is checked first regardless, so a typo fails on the one-comparison report a caller develops against instead of waiting for the second metric.
Raises:
ValueError: if method is not a known correction, or if any comparison
already carries an (adj:...) suffix. Correcting twice compounds
the p-values into a number that means nothing, and merged reports
make a second pass easy to reach.
analyze¶
analyze(scores: dict[str, np.ndarray], score_types: str | dict[str, str] | None = None, rubrics: dict[str, MetricRubric] | None = None, calibration_data: dict | None = None, comparison_scores: dict[str, np.ndarray] | None = None, gates: dict[str, float | GateSpec] | None = None, metric_directions: dict[str, bool] | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, adjust_method: AdjustMethod | None = None) -> StatisticalReport
Orchestrate a full statistical analysis pipeline.
Args:
scores: Dict mapping metric name to 1D score arrays.
score_types: Score type ("binary", "ordinal", "continuous") applied
to every metric when given as a bare string, or a dict mapping
metric name to score type. Auto-detected if absent.
rubrics: Optional dict mapping metric name to MetricRubric
(used for ordinal distribution).
calibration_data: Optional dict with keys "judge", "human" mapping
to arrays for PPI bias correction.
comparison_scores: Optional dict mapping metric name to 1D score
arrays from a second system (for paired comparison).
gates: Optional dict mapping metric name to threshold values — a
bare float (gated at the default lower_ci strictness) or a
:class:~latent.gates.gating.GateSpec carrying its own
strictness (e.g. point_estimate for small-n binary rates).
metric_directions: Optional dict mapping metric name to a boolean
indicating lower_is_better. When None (the default),
every metric is treated as higher-is-better. A GateSpec
in gates that pins its own lower_is_better wins for that
metric's gate and its comparison; pinning the opposite of the
direction given here raises ValueError.
n_resamples: Number of bootstrap resamples.
confidence_level: Confidence level for intervals.
seed: Random seed for reproducibility.
adjust_method: Optional family-wise correction for the comparison
p-values. "holm" controls FWER (recommended default when
you opt in), "fdr_bh" controls FDR. None (default)
leaves p-values uncorrected — callers comparing more than one
metric on the same eval set should opt in to avoid alpha
inflation. Only the comparison p-values are adjusted; the
single-system metric CIs are unaffected.
Returns: StatisticalReport populated with metrics, comparisons, gates, ordinal distributions, and method explanations.
explain_method¶
explain_method(metric_name: str, score_type: str, sample_size: int, has_calibration: bool = False, is_comparison: bool = False) -> str
Explain which statistical method was used for a metric and why.
Args: metric_name: Name of the metric being analyzed. score_type: One of "binary", "ordinal", "continuous". sample_size: Number of samples. has_calibration: Whether calibration data was available. is_comparison: Whether a system comparison was performed.
Returns: Human-readable explanation of the method used and the reasoning.