latent.flows.retrieval_eval_flow¶
retrieval_eval_flow — Evaluate retrieval quality before full agent evaluation.
Functions¶
retrieval_eval_flow¶
retrieval_eval_flow(eval_data: pd.DataFrame, retriever: Any, question_column: str = 'question', expected_column: str = 'expected', id_column: str | None = None, k: int = 10, judge_model: str = 'claude-sonnet-4-5-20250929', answerability_check: bool = True, chunk_relevance_labels: bool = True, coverage_grading: bool = True, n_resamples: int = 10000, confidence_level: float = 0.95, seed: int | None = None, summary_model: str | None = None, summary_label: str | None = None) -> dict[str, Any]
Evaluate retrieval quality before running full agent eval.
For each question+expected answer pair, retrieves chunks and optionally: 1. Classifies whether the question needs retrieval at all. 2. Grades coverage (relevance 1-5, info completeness 1-5). 3. Labels each chunk as relevant/contradicts/not_relevant.
Args: eval_data: DataFrame with question and expected answer columns. retriever: Object with .search(query, k) method returning chunks. question_column: Column with questions. expected_column: Column with expected/ground truth answers. id_column: Optional ID column. k: Number of chunks to retrieve. judge_model: LLM model for evaluation. answerability_check: Whether to classify answerability. chunk_relevance_labels: Whether to label individual chunks. coverage_grading: Whether to grade coverage. n_resamples: Bootstrap resamples. confidence_level: CI confidence level. seed: Random seed.
Returns: Dict with: - results: list[dict] per-row retrieval quality labels. - report: StatisticalReport with coverage metrics. - markdown: str rendered report. - relevant_chunk_ids: dict[row_id, list[chunk_id]] mapping.