Skip to content

latent.stats.report

Report orchestrator and method explanation (FR-8.1 + FR-8.4).

Functions

adjust_comparisons

adjust_comparisons(comparisons: list[ComparisonResult], method: AdjustMethod) -> list[ComparisonResult]

Apply a family-wise correction to a set of comparison p-values.

The one place the corrected p-value, the (adj:<method>) method suffix and the rewritten interpretation tail are produced together. Every builder (compare_systems and friends) embeds the verdict and raw p-value at the tail of interpretation (e.g. "(statistically significant, p=0.0300)"); after correction the p-value changes, sometimes crossing the significance threshold, so the text is patched to match. Without it callers read p_value=0.09 beside "...significant, p=0.0300".

A single comparison is returned untouched — there is no family to correct. method is checked first regardless, so a typo fails on the one-comparison report a caller develops against instead of waiting for the second metric.

Raises: ValueError: if method is not a known correction, or if any comparison already carries an (adj:...) suffix. Correcting twice compounds the p-values into a number that means nothing, and merged reports make a second pass easy to reach.

analyze

analyze(scores: dict[str, np.ndarray], score_types: str | dict[str, str] | None = None, rubrics: dict[str, MetricRubric] | None = None, calibration_data: dict | None = None, comparison_scores: dict[str, np.ndarray] | None = None, gates: dict[str, float | GateSpec] | None = None, metric_directions: dict[str, bool] | None = None, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, adjust_method: AdjustMethod | None = None) -> StatisticalReport

Orchestrate a full statistical analysis pipeline.

Args: scores: Dict mapping metric name to 1D score arrays. score_types: Score type ("binary", "ordinal", "continuous") applied to every metric when given as a bare string, or a dict mapping metric name to score type. Auto-detected if absent. rubrics: Optional dict mapping metric name to MetricRubric (used for ordinal distribution). calibration_data: Optional dict with keys "judge", "human" mapping to arrays for PPI bias correction. comparison_scores: Optional dict mapping metric name to 1D score arrays from a second system (for paired comparison). gates: Optional dict mapping metric name to threshold values — a bare float (gated at the default lower_ci strictness) or a :class:~latent.gates.gating.GateSpec carrying its own strictness (e.g. point_estimate for small-n binary rates). metric_directions: Optional dict mapping metric name to a boolean indicating lower_is_better. When None (the default), every metric is treated as higher-is-better. A GateSpec in gates that pins its own lower_is_better wins for that metric's gate and its comparison; pinning the opposite of the direction given here raises ValueError. n_resamples: Number of bootstrap resamples. confidence_level: Confidence level for intervals. seed: Random seed for reproducibility. adjust_method: Optional family-wise correction for the comparison p-values. "holm" controls FWER (recommended default when you opt in), "fdr_bh" controls FDR. None (default) leaves p-values uncorrected — callers comparing more than one metric on the same eval set should opt in to avoid alpha inflation. Only the comparison p-values are adjusted; the single-system metric CIs are unaffected.

Returns: StatisticalReport populated with metrics, comparisons, gates, ordinal distributions, and method explanations.

explain_method

explain_method(metric_name: str, score_type: str, sample_size: int, has_calibration: bool = False, is_comparison: bool = False) -> str

Explain which statistical method was used for a metric and why.

Args: metric_name: Name of the metric being analyzed. score_type: One of "binary", "ordinal", "continuous". sample_size: Number of samples. has_calibration: Whether calibration data was available. is_comparison: Whether a system comparison was performed.

Returns: Human-readable explanation of the method used and the reasoning.

Attributes

AdjustMethod