Skip to content

latent.agents.eval.judge

Judge[T] -- evaluation scoring agent (FR-1.4).

Classes

Judge

Judge(name: str, model: str, output_type: type[T], prompt_template: str | None = None, system_prompt: str | None = None, few_shot: Sequence[Message] | None = None, temperature: float = 0.0, max_tokens: int = 4096, is_eligible: Callable[[dict], bool] | None = None, row_mapper: Callable[[dict], dict] | None = None, column_prefix: str | None = None, kwargs: Any = {})

LLM agent specialized for scoring data rows.

Inherits ReActAgent for LLM access and adds: - output_type: Pydantic model defining the score schema + rubrics - prompt_template: Format string for constructing prompts from row data. If omitted, auto-generated from output_type's score annotations. - few_shot: Worked examples as alternating user/assistant messages, replayed ahead of every scored row - evaluate(row) -> T: Builds messages, calls run(), parses output - call(row) -> T: Delegates to evaluate() for @task.map() compat

Scoring is async-only: evaluate and __call__ are coroutines, so a synchronous call site adopts a Judge by restructuring around await (or asyncio.run), not by swapping one call for another.

Methods

Judge.evaluate

evaluate(row: dict) -> T

Score a single data row.

Applies row_mapper to translate the raw row into the keys prompt_template expects, formats the prompt, calls run() with structured output, and parses the JSON response.

few_shot examples precede the scored row, after the system message, so every call sees the same conditioning.

Attributes

DEFAULT_SYSTEM_PROMPT

RUBRICLESS_PROMPT_TEMPLATE