Skip to content

latent.flows.conversation_simulation.flow

conversation_simulation_flow — Prefect-orchestrated simulation + scoring.

Classes

JudgeScoringConfig

JudgeScoringConfig()

Configuration for judge scoring behavior.

SimulationGuardrails

SimulationGuardrails()

Guardrails for simulation behavior.

Functions

conversation_simulation_flow

conversation_simulation_flow(agent_class: type, agent_kwargs: dict[str, Any] | None = None, scenarios: pd.DataFrame | None = None, human_class: type | None = None, human_kwargs: dict[str, Any] | None = None, judge_class: Callable[[], Judge[Any]] | None = None, judge_scoring: JudgeScoringConfig | None = None, simulation_guardrails: SimulationGuardrails | None = None, context_column: str = 'context') -> dict[str, Any]

Simulate conversations and optionally score them.

Args: agent_class: Agent under test class. agent_kwargs: Constructor kwargs for the agent. scenarios: Scenario data. Loaded from catalog if not provided. human_class: Human agent class (defaults to HumanAgent). human_kwargs: Constructor kwargs for the human agent. judge_class: Judge factory callable. If set, scoring is enabled. Gates are auto-derived from the judge's output_type pass_threshold annotations. judge_scoring: Scoring config dict. Keys: - scope (str): "conversation" (default) scores each conversation as a whole; "turn" scores each assistant turn then aggregates. - aggregation (str): Turn aggregation strategy ('mean', 'min', 'max', 'last'). Only for turn scope. - gates (dict): Optional quality gates overriding auto-derived thresholds. simulation_guardrails: Guardrails dict. Keys: - max_turns_per_conversation (int): Max turns per conversation (default 30). context_column: Column name for scenario context (default "context").

Returns: Dict with 'simulation_results' (DataFrame) and 'report' (scoring report dict or None).

results_to_conversations

results_to_conversations(results: pd.DataFrame) -> list[Conversation]

Convert simulation results DataFrame to Conversation objects.

Filters to successful simulations and converts turn dicts to Conversation objects with typed Turn entries.

Args: results: DataFrame from simulation with 'success' and 'conversation' columns.

Returns: List of Conversation objects for successful simulations.