Skip to content

latent.flows.guardrail_redteam.flow

Guardrail red-team evaluation flow.

Unit mode: tests prompts against input scanners directly (no LLM calls). E2E mode: routes prompts through a full agent + GuardrailMiddleware.

Functions

guardrail_redteam

guardrail_redteam(inputs: list[dict], agent: object | None = None, pre_scanners: list[object] | None = None, post_scanners: list[object] | None = None) -> list[dict]

Run a guardrail red-team evaluation against a list of test cases.

Selects unit mode (scanner-only, no LLM calls) when agent is None, and E2E mode when an agent is provided.

Args: inputs: list of {"prompt": str, "expected_blocked": bool, "label": str} agent: optional agent with a stream() method; omit for unit mode pre_scanners: list of input scanner instances post_scanners: optional list of output scanner instances (E2E only)

Returns: list of result dicts — see :func:run_unit_eval / :func:run_e2e_eval

Example::

from latent.flows.guardrail_redteam.flow import guardrail_redteam
from latent.guardrails.scanners.builtin import LanguageScanner

results = guardrail_redteam(
    inputs=[
        {"prompt": "hello", "expected_blocked": False, "label": "en"},
        {"prompt": "bonjour", "expected_blocked": True, "label": "fr"},
    ],
    pre_scanners=[LanguageScanner(allowed_languages=["en"])],
)
for r in results:
    print(r["label"], r["blocked"], r["score"])

run_e2e_eval

run_e2e_eval(test_cases: list[dict], agent: object, pre_scanners: list[object], post_scanners: list[object] | None = None) -> list[dict]

E2E mode: routes prompts through a full agent wrapped with GuardrailMiddleware.

Args: test_cases: list of {"prompt": str, "expected_blocked": bool, "label": str} agent: any agent with a stream() method pre_scanners: list of input scanner instances post_scanners: optional list of output scanner instances

Returns: list of result dicts with blocked, violated_rules, response, score fields

run_unit_eval

run_unit_eval(test_cases: list[dict], pre_scanners: list[object]) -> list[dict]

Unit mode: test each prompt against pre scanners only (no agent call).

Args: test_cases: list of {"prompt": str, "expected_blocked": bool, "label": str} pre_scanners: list of scanner instances with scan(prompt) method

Returns: list of result dicts with blocked, violated_rules, score fields