Skip to content

latent.stats.drift

Drift detection between evaluation runs.

Functions

detect_drift

detect_drift(baseline_scores: np.ndarray, current_scores: np.ndarray, metric_name: str = '', score_type: str = 'binary', confidence_level: float = 0.95, seed: int | None = None) -> DriftResult

Compare two eval runs to detect performance drift.

For binary scores, uses bootstrap CI on the difference of means. For ordinal scores, uses the Mann-Whitney U test via scipy.

Args: baseline_scores: Scores from the baseline eval run. current_scores: Scores from the current eval run. metric_name: Name of the metric being compared. score_type: Type of scores - "binary" or "ordinal". confidence_level: Confidence level for CIs. seed: Random seed for reproducibility.

Returns: DriftResult with severity, delta, CI, p-value, and effect size.

drift_report

drift_report(baseline_scores: dict[str, np.ndarray], current_scores: dict[str, np.ndarray], score_types: dict[str, str] | None = None, confidence_level: float = 0.95, seed: int | None = None) -> list[DriftResult]

Run drift detection across multiple metrics.

Args: baseline_scores: Dict mapping metric name to baseline score arrays. current_scores: Dict mapping metric name to current score arrays. score_types: Optional dict mapping metric name to score type ("binary" or "ordinal"). Defaults to "binary" for all. confidence_level: Confidence level for CIs. seed: Random seed for reproducibility.

Returns: List of DriftResult, sorted by severity (most severe first).

multi_run_trend

multi_run_trend(runs: list[np.ndarray], metric_name: str = '', confidence_level: float = 0.95, seed: int | None = None, higher_is_better: bool = True) -> dict

Track a metric across 3+ runs to detect trends.

Args: runs: List of score arrays, one per run (ordered chronologically). metric_name: Name of the metric. confidence_level: Confidence level for CIs. seed: Random seed for reproducibility. higher_is_better: If True (default), increasing values are "improving". Set to False for metrics where lower is better (e.g. latency, error rate).

Returns: Dict with keys: - "values": list of mean scores per run - "trend": "improving", "degrading", or "stable" - "is_monotonic": bool - whether values are strictly monotonic - "cis": list of (lower, upper) tuples per run