Skip to content

latent.optimize.autoresearch.optimizer

Autonomous research optimizer using the Claude Agent SDK.

Classes

AutoResearchOptimizer

AutoResearchOptimizer(brief: ResearchBrief, eval_fn: Callable[[list[dict[str, Any]]], Awaitable[dict[str, float]]], dataset: pd.DataFrame | None = None, stratify_column: str | None = None, sample_size: int | None = None, confirmation_size: int | None = None, min_per_stratum: int = 2, repo_root: Path = Path('.'), max_iterations: int = 20, patience: int = 5, checks: list[str] | None = None, agent_model: str = 'claude-opus-4-6', agent_max_turns: int = 15, allowed_tools: list[str] | None = None, tracker: ExperimentTracker | None = None, checkpoint_path: Path | None = None, check_timeout: int = 300, eval_timeout: int = 600, secondary_metrics: list[str] | None = None, system_prompt: str | None = None, scope_paths: list[str] | None = None, validation_split: float = 0.0)

Autonomous code optimizer driven by the Claude Agent SDK.

Each iteration: 1. Runs a Claude Agent SDK session — agent reads codebase, applies one change. 2. Runs quality checks (pytest subset) — controlled by the optimizer, not the agent. 3. Awaits eval_fn(sample) — caller-supplied async eval on a dataset sample. 4. Optionally confirms on a larger sample (progressive evaluation). 5. Commits (keep) or hard-resets (discard) via git. 6. Tracks every experiment via ExperimentTracker.

Features: - Stratified subsampling with 3-tier progressive evaluation. - Experiment insights — classifies results as worked/failed/promising and injects analysis into the next iteration's prompt. - Failure analysis — per-category score breakdown injected into prompt. - Progressive difficulty — oversamples from historically weak categories. - CI-aware decisions — uses lower confidence bound for keep/discard. - Rollback on deep regression — reverts to best if consecutive deep drops.

Methods

AutoResearchOptimizer.optimize

optimize() -> OptimizationResult

Run the autonomous research loop.

Attributes

DEFAULT_ALLOWED_TOOLS