Skip to content

latent.stats.conversation

Conversation schema and turn-position analysis (FR-5.1 + FR-5.4).

Classes

Conversation

Conversation()

A multi-turn conversation.

Turn

Turn()

A single turn in a conversation.

Functions

aggregate_turn_scores

aggregate_turn_scores(scores: np.ndarray, strategy: str = 'mean', weights: np.ndarray | None = None) -> float

Aggregate per-turn scores with a configurable strategy.

Args: scores: 1-D array of per-turn scores. strategy: One of "mean", "min", "max", "weighted", "last". weights: Required when strategy is "weighted"; same length as scores.

Returns: Single aggregate float.

Raises: ValueError: If strategy is unknown or weights are missing/mismatched.

turn_level_scores

turn_level_scores(conversation: Conversation, score_fn: Callable[[Turn], float]) -> np.ndarray

Apply a scoring function to each turn, returning per-turn scores.

Args: conversation: The conversation to score. score_fn: Function that maps a Turn to a float score.

Returns: 1-D array of per-turn scores.

turn_position_analysis

turn_position_analysis(conversations: list[Conversation], score_fn: Callable[[Turn], float], max_turns: int | None = None, confidence_level: float = 0.95, seed: int | None = None) -> list[MetricResult]

Analyse scores by turn position across multiple conversations.

For each turn position, aggregate scores across all conversations that have at least that many turns, then compute a bootstrap CI.

Args: conversations: List of conversations to analyse. score_fn: Function that maps a Turn to a float score. max_turns: If set, only analyse up to this many turn positions. confidence_level: Confidence level for bootstrap CIs. seed: Random seed for reproducibility.

Returns: List of MetricResult, one per turn position.