latent.mlflow.evaluate¶
MLflow evaluation wrapper for LLM agent evaluation workflows.
Functions¶
evaluate¶
evaluate(data: pd.DataFrame, agent: Callable | Any, evaluators: list[str | Callable] | str = 'default', model_type: str = 'text', targets: str | None = None, predictions: str | None = None, kwargs = {}) -> pd.DataFrame
Evaluate an agent/model with automatic MLflow tracking.
This is a thin wrapper around mlflow.evaluate() that: - Automatically infers column mappings from catalog.yaml - Handles both string evaluator names and custom evaluators - Logs results to the current flow's MLflow run - Returns results as a DataFrame for further processing
Args: data: Evaluation dataset (DataFrame) agent: Agent function or callable object to evaluate evaluators: List of evaluators (built-in names or custom functions) model_type: Type of model ("text", "question-answering", etc.) targets: Column name for ground truth (optional) predictions: Column name for predictions (optional) **kwargs: Additional arguments passed to mlflow.evaluate()
Returns: DataFrame with evaluation results (metrics)
Example: @task("evaluate", input="eval_data", output="results") def evaluate_task(eval_data: pd.DataFrame): agent = Agent(model=params.model) results = evaluate( data=eval_data, agent=agent, evaluators=params.evaluators ) return results
evaluator¶
Decorator to create custom MLflow evaluators.
Wraps mlflow.metrics.make_metric() with a cleaner API.
Args: name: Name of the metric (defaults to function name) greater_is_better: Whether higher values are better **metric_kwargs: Additional arguments for make_metric()
Example: @evaluator(name="custom_score") def custom_evaluator(eval_df, builtin_metrics): # Access predictions and ground truth predictions = eval_df["outputs"] ground_truth = eval_df["ground_truth"]
# Compute your metric
scores = []
for pred, truth in zip(predictions, ground_truth):
score = compute_similarity(pred, truth)
scores.append(score)
# Return as list (wrapper handles MetricValue creation)
return scores
Example with LLM judge: @evaluator(name="llm_judge_score") def llm_judge(eval_df, builtin_metrics): from litellm import completion
scores = []
for _, row in eval_df.iterrows():
prompt = f"Rate this response: {row['outputs']}"
response = completion(
model="gpt-4",
messages=[{"role": "user", "content": prompt}]
)
score = parse_score(response.choices[0].message.content)
scores.append(score)
return scores