Skip to content

latent.stats.categories

Category-level statistical aggregation and record builders.

Functions

build_failure_mode_summary

build_failure_mode_summary(records: list[RecordResult], taxonomy: dict[str, str] | None = None) -> list[FailureModeSummary]

Count failure modes and compute proportions.

Args: records: Records that may have .failure_mode set. taxonomy: Optional {"mode": "description"} dict for descriptions.

Returns: List of FailureModeSummary sorted by count descending.

build_records_from_df

build_records_from_df(scored_data: pd.DataFrame, score_columns: list[str], question_column: str = 'question', output_column: str = 'output', expected_column: str | None = None, category_column: str | None = None, id_column: str | None = None, tag_column: str | None = None, failure_mode_column: str | None = None) -> list[RecordResult]

Convert a scored DataFrame to a list of RecordResult objects.

Args: scored_data: DataFrame with evaluation results. score_columns: Column names to include as scores. question_column: Column with the input question. output_column: Column with the agent output. expected_column: Column with the expected/ground truth answer. category_column: Column with the category label. id_column: Column with the record ID. tag_column: Column with optional tag. failure_mode_column: Column with failure mode classification.

Returns: List of RecordResult objects.

category_breakdown

category_breakdown(records: list[RecordResult], score_name: str, score_type: str = 'binary', confidence_level: float = 0.95, n_resamples: int = 10000, seed: int | None = None) -> list[CategorySummary]

Aggregate a named score per category with confidence intervals.

Groups records by .category, computes a MetricResult per group. Binary scores use Wilson CI; continuous/ordinal use bootstrap CI.

Args: records: Evaluation records with .category and .scores fields. score_name: Key in record.scores to aggregate. score_type: "binary", "continuous", or "ordinal". confidence_level: CI confidence level. n_resamples: Bootstrap resamples (for non-binary). seed: Random seed.

Returns: List of CategorySummary, sorted by category name.

merge_category_summaries

merge_category_summaries(summaries: list[CategorySummary]) -> list[CategorySummary]

Merge CategorySummary objects with the same category name.

When category_breakdown() is called once per score column, it produces separate summaries per category. This function merges them into one summary per category with all metrics combined.