Skip to content

latent.flows.autoresearch_flow

Prebuilt AutoResearch flow -- autonomous code optimization with eval subflow.

Handles stratified subsampling, progressive confirmation, and the optimizer loop. The caller only needs to supply a Prefect @flow that evaluates a sample and returns a :class:~latent.stats.models.StatisticalReport.

When deployment_name is provided, each eval iteration is submitted to a Prefect work pool via run_deployment, giving full process isolation (no aiohttp event loop leaks between the Claude Agent SDK and litellm).

Usage (consumer side)::

from latent.prefect import flow
from latent.stats.models import StatisticalReport

@flow("my_eval")
def my_eval(sample_pairs: list[dict]) -> StatisticalReport:
    inference = agent_inference_flow(sample_pairs=sample_pairs)
    report = agent_judge_flow(inference_results=inference)
    return report["statistical_report_obj"]

autoresearch_agent_flow(
    eval_flow=my_eval,
    brief=brief,
    dataset=df,
    deployment_name="autoresearch_eval/autoresearch-eval",
)

Functions

autoresearch_agent_flow

autoresearch_agent_flow(eval_flow: Any, brief: ResearchBrief, dataset: pd.DataFrame, deployment_name: str | None = None, stratify_column: str = 'category', sample_size: int = 60, confirmation_size: int | None = 150, min_per_stratum: int = 2, repo_root: Path = Path('.'), max_iterations: int = 20, patience: int = 5, checks: list[str] | None = None, agent_model: str = 'claude-opus-4-6', agent_max_turns: int = 15, allowed_tools: list[str] | None = None, tracker: ExperimentTracker | None = None, checkpoint_path: Path | None = None, check_timeout: int = 300, eval_timeout: int = 600, secondary_metrics: list[str] | None = None, system_prompt: str | None = None, scope_paths: list[str] | None = None) -> dict[str, Any]

Run autonomous code optimization with a user-defined eval subflow.

The eval_flow must be a Prefect @flow-decorated function with the signature (sample_pairs: list[dict]) -> StatisticalReport.

When deployment_name is provided (e.g. "autoresearch_eval/autoresearch-eval"), each eval is submitted to a Prefect work pool via run_deployment for full process isolation. Otherwise falls back to run_flow_sync in a thread.