Skip to content

latent.flows.ner_flow

ner_flow — Evaluate named entity recognition quality.

Functions

ner_flow

ner_flow(eval_data: pd.DataFrame, predicted_column: str = 'predicted_spans', gold_column: str = 'gold_spans', match_mode: MatchMode = 'exact', iou_threshold: float = 0.5, labels: list[str] | None = None, gates: dict[str, float] | None = None, confidence_level: float = 0.95, n_resamples: int = 10000, seed: int | None = None) -> dict[str, Any]

Evaluate NER with span-level precision, recall, F1.

The input DataFrame must have columns containing lists of span dicts with keys: start, end, label.

Args: eval_data: DataFrame with span columns. predicted_column: Column with predicted spans (list of dicts). gold_column: Column with gold spans (list of dicts). match_mode: "exact", "overlap", or "iou". iou_threshold: IoU threshold for match_mode="iou". labels: Entity types to report on. gates: Quality thresholds for metrics. confidence_level: CI confidence level. n_resamples: Bootstrap resamples. seed: Random seed.

Returns: Dict with: metrics, all_passed, report.