latent.flows.conversation_simulation.flow¶
conversation_simulation_flow — Prefect-orchestrated simulation + scoring.
Classes¶
JudgeScoringConfig¶
Configuration for judge scoring behavior.
SimulationGuardrails¶
Guardrails for simulation behavior.
Functions¶
conversation_simulation_flow¶
conversation_simulation_flow(agent_class: type, agent_kwargs: dict[str, Any] | None = None, scenarios: pd.DataFrame | None = None, human_class: type | None = None, human_kwargs: dict[str, Any] | None = None, judge_class: Callable[[], Judge[Any]] | None = None, judge_scoring: JudgeScoringConfig | None = None, simulation_guardrails: SimulationGuardrails | None = None, context_column: str = 'context') -> dict[str, Any]
Simulate conversations and optionally score them.
Args:
agent_class: Agent under test class.
agent_kwargs: Constructor kwargs for the agent.
scenarios: Scenario data. Loaded from catalog if not provided.
human_class: Human agent class (defaults to HumanAgent).
human_kwargs: Constructor kwargs for the human agent.
judge_class: Judge factory callable. If set, scoring
is enabled. Gates are auto-derived from the judge's
output_type pass_threshold annotations.
judge_scoring: Scoring config dict. Keys:
- scope (str): "conversation" (default) scores
each conversation as a whole; "turn" scores each
assistant turn then aggregates.
- aggregation (str): Turn aggregation strategy
('mean', 'min', 'max', 'last'). Only for turn scope.
- gates (dict): Optional quality gates overriding
auto-derived thresholds.
simulation_guardrails: Guardrails dict. Keys:
- max_turns_per_conversation (int): Max turns per
conversation (default 30).
context_column: Column name for scenario context (default "context").
Returns: Dict with 'simulation_results' (DataFrame) and 'report' (scoring report dict or None).
results_to_conversations¶
Convert simulation results DataFrame to Conversation objects.
Filters to successful simulations and converts turn dicts to Conversation objects with typed Turn entries.
Args: results: DataFrame from simulation with 'success' and 'conversation' columns.
Returns: List of Conversation objects for successful simulations.