latent.flows.rag_research_flow.flow¶
rag_research_flow — RAG pipeline optimization with flow+task pattern.
Functions¶
rag_research_flow¶
rag_research_flow(documents: list[str] | None = None, eval_data: list[dict[str, Any]] | None = None, eval_fn: RAGEvalFn | None = None, metrics: list | None = None, baseline_config: dict[str, Any] | None = None, max_trials: int | None = None, max_eval_configs: int | None = None, embedding_provider: str | None = None, embedding_model: str | None = None, persist_base: str | None = None, chunk_sizes: list[int] | None = None, chunk_overlaps: list[int] | None = None, backends: list[str] | None = None, alphas: list[float] | None = None) -> dict[str, Any]
Run RAG pipeline optimization with MLflow logging.
Two-phase flow+task pattern:
1. search — Optuna Bayesian optimization over RAG component space.
2. eval — Validate top configs via caller-supplied eval_fn.
Args:
documents: Raw texts to index (required, non-empty).
eval_data: Evaluation dataset with "query" key (required, non-empty).
eval_fn: Async callable (config, sample) -> EvalReport.
If None, eval phase is skipped.
metrics: RAGMetric instances. Defaults to [ContentRecallMetric(k=7), ContentMRRMetric()].
baseline_config: Baseline pipeline config for Optuna seeding.
max_trials: Search phase trial budget.
max_eval_configs: Number of top configs to validate in eval phase.
embedding_provider: Embedding provider (e.g. "voyage").
embedding_model: Embedding model name.
persist_base: Directory for caching pre-indexed adapters.
chunk_sizes: Chunk sizes to explore.
chunk_overlaps: Chunk overlaps to explore.
backends: Backend types to explore.
alphas: Hybrid alpha values to explore.
Returns:
Dict with report (StatisticalReport), markdown (str),
optimization (OptimizationReport), and scalar result fields.
search_configs¶
search_configs(documents: list[str], eval_data: list[dict[str, Any]], metrics: list, max_trials: int = 15, baseline_config: dict[str, Any] | None = None, persist_base: str | None = None, embedding_provider: str = 'voyage', embedding_model: str = 'voyage-3', chunk_sizes: list[int] | None = None, chunk_overlaps: list[int] | None = None, backends: list[str] | None = None, alphas: list[float] | None = None) -> list[Experiment]
Search phase: Optuna Bayesian optimization over RAG component space.