Skip to content

latent.flows.rag_research_flow.flow

rag_research_flow — RAG pipeline optimization with flow+task pattern.

Functions

rag_research_flow

rag_research_flow(documents: list[str] | None = None, eval_data: list[dict[str, Any]] | None = None, eval_fn: RAGEvalFn | None = None, metrics: list | None = None, baseline_config: dict[str, Any] | None = None, max_trials: int | None = None, max_eval_configs: int | None = None, embedding_provider: str | None = None, embedding_model: str | None = None, persist_base: str | None = None, chunk_sizes: list[int] | None = None, chunk_overlaps: list[int] | None = None, backends: list[str] | None = None, alphas: list[float] | None = None) -> dict[str, Any]

Run RAG pipeline optimization with MLflow logging.

Two-phase flow+task pattern: 1. search — Optuna Bayesian optimization over RAG component space. 2. eval — Validate top configs via caller-supplied eval_fn.

Args: documents: Raw texts to index (required, non-empty). eval_data: Evaluation dataset with "query" key (required, non-empty). eval_fn: Async callable (config, sample) -> EvalReport. If None, eval phase is skipped. metrics: RAGMetric instances. Defaults to [ContentRecallMetric(k=7), ContentMRRMetric()]. baseline_config: Baseline pipeline config for Optuna seeding. max_trials: Search phase trial budget. max_eval_configs: Number of top configs to validate in eval phase. embedding_provider: Embedding provider (e.g. "voyage"). embedding_model: Embedding model name. persist_base: Directory for caching pre-indexed adapters. chunk_sizes: Chunk sizes to explore. chunk_overlaps: Chunk overlaps to explore. backends: Backend types to explore. alphas: Hybrid alpha values to explore.

Returns: Dict with report (StatisticalReport), markdown (str), optimization (OptimizationReport), and scalar result fields.

search_configs

search_configs(documents: list[str], eval_data: list[dict[str, Any]], metrics: list, max_trials: int = 15, baseline_config: dict[str, Any] | None = None, persist_base: str | None = None, embedding_provider: str = 'voyage', embedding_model: str = 'voyage-3', chunk_sizes: list[int] | None = None, chunk_overlaps: list[int] | None = None, backends: list[str] | None = None, alphas: list[float] | None = None) -> list[Experiment]

Search phase: Optuna Bayesian optimization over RAG component space.