Skip to content

latent.stats.structured

Structured output evaluation metrics.

Functions

composite_accuracy

composite_accuracy(outputs: list[dict], expected: list[dict], fields: list[str] | None = None) -> MetricResult

Percentage of fields correct per example, aggregated with bootstrap CI.

For each example, computes the fraction of fields that match exactly. Aggregates these fractions with bootstrap CI.

Args: outputs: List of actual output dicts. expected: List of expected output dicts (same length as outputs). fields: Which fields to evaluate. If None, uses union of all keys in expected.

Returns: MetricResult with mean composite accuracy and bootstrap CI.

field_level_accuracy

field_level_accuracy(outputs: list[dict], expected: list[dict], fields: list[str] | None = None) -> dict[str, MetricResult]

Per-field exact match accuracy with Wilson CIs.

Args: outputs: List of actual output dicts. expected: List of expected output dicts (same length as outputs). fields: Which fields to evaluate. If None, uses union of all keys.

Returns: Dict mapping field name -> MetricResult with Wilson CI.

result_set_match

result_set_match(actual_sets: list[list[dict]], expected_sets: list[list[dict]]) -> MetricResult

Row-level order-independent exact match between result sets.

For each pair of (actual, expected), checks whether the sets contain the same rows regardless of order. Returns Wilson CI on match rate.

Args: actual_sets: List of actual result sets (each a list of dicts). expected_sets: List of expected result sets (same length).

Returns: MetricResult with match rate and Wilson CI.

schema_compliance

schema_compliance(outputs: list[dict], schema: dict) -> MetricResult

Binary pass/fail for each output against a JSON schema.

Args: outputs: List of output dicts to validate. schema: JSON Schema dict to validate against.

Returns: MetricResult with compliance rate and Wilson CI.