latent.stats.structured¶
Structured output evaluation metrics.
Functions¶
composite_accuracy¶
composite_accuracy(outputs: list[dict], expected: list[dict], fields: list[str] | None = None) -> MetricResult
Percentage of fields correct per example, aggregated with bootstrap CI.
For each example, computes the fraction of fields that match exactly. Aggregates these fractions with bootstrap CI.
Args: outputs: List of actual output dicts. expected: List of expected output dicts (same length as outputs). fields: Which fields to evaluate. If None, uses union of all keys in expected.
Returns: MetricResult with mean composite accuracy and bootstrap CI.
field_level_accuracy¶
field_level_accuracy(outputs: list[dict], expected: list[dict], fields: list[str] | None = None) -> dict[str, MetricResult]
Per-field exact match accuracy with Wilson CIs.
Args: outputs: List of actual output dicts. expected: List of expected output dicts (same length as outputs). fields: Which fields to evaluate. If None, uses union of all keys.
Returns: Dict mapping field name -> MetricResult with Wilson CI.
result_set_match¶
Row-level order-independent exact match between result sets.
For each pair of (actual, expected), checks whether the sets contain the same rows regardless of order. Returns Wilson CI on match rate.
Args: actual_sets: List of actual result sets (each a list of dicts). expected_sets: List of expected result sets (same length).
Returns: MetricResult with match rate and Wilson CI.
schema_compliance¶
Binary pass/fail for each output against a JSON schema.
Args: outputs: List of output dicts to validate. schema: JSON Schema dict to validate against.
Returns: MetricResult with compliance rate and Wilson CI.