Skip to content

latent.mlflow.evaluate

MLflow evaluation wrapper for LLM agent evaluation workflows.

Functions

evaluate

evaluate(data: pd.DataFrame, agent: Callable | Any, evaluators: list[str | Callable] | str = 'default', model_type: str = 'text', targets: str | None = None, predictions: str | None = None, kwargs = {}) -> pd.DataFrame

Evaluate an agent/model with automatic MLflow tracking.

This is a thin wrapper around mlflow.evaluate() that: - Automatically infers column mappings from catalog.yaml - Handles both string evaluator names and custom evaluators - Logs results to the current flow's MLflow run - Returns results as a DataFrame for further processing

Args: data: Evaluation dataset (DataFrame) agent: Agent function or callable object to evaluate evaluators: List of evaluators (built-in names or custom functions) model_type: Type of model ("text", "question-answering", etc.) targets: Column name for ground truth (optional) predictions: Column name for predictions (optional) **kwargs: Additional arguments passed to mlflow.evaluate()

Returns: DataFrame with evaluation results (metrics)

Example: @task("evaluate", input="eval_data", output="results") def evaluate_task(eval_data: pd.DataFrame): agent = Agent(model=params.model) results = evaluate( data=eval_data, agent=agent, evaluators=params.evaluators ) return results

evaluator

evaluator(name: str | None = None, greater_is_better: bool = True, metric_kwargs = {})

Decorator to create custom MLflow evaluators.

Wraps mlflow.metrics.make_metric() with a cleaner API.

Args: name: Name of the metric (defaults to function name) greater_is_better: Whether higher values are better **metric_kwargs: Additional arguments for make_metric()

Example: @evaluator(name="custom_score") def custom_evaluator(eval_df, builtin_metrics): # Access predictions and ground truth predictions = eval_df["outputs"] ground_truth = eval_df["ground_truth"]

    # Compute your metric
    scores = []
    for pred, truth in zip(predictions, ground_truth):
        score = compute_similarity(pred, truth)
        scores.append(score)

    # Return as list (wrapper handles MetricValue creation)
    return scores

Example with LLM judge: @evaluator(name="llm_judge_score") def llm_judge(eval_df, builtin_metrics): from litellm import completion

    scores = []
    for _, row in eval_df.iterrows():
        prompt = f"Rate this response: {row['outputs']}"
        response = completion(
            model="gpt-4",
            messages=[{"role": "user", "content": prompt}]
        )
        score = parse_score(response.choices[0].message.content)
        scores.append(score)

    return scores