latent.flows.guardrail_redteam.flow¶
Guardrail red-team evaluation flow.
Unit mode: tests prompts against input scanners directly (no LLM calls). E2E mode: routes prompts through a full agent + GuardrailMiddleware.
Functions¶
guardrail_redteam¶
guardrail_redteam(inputs: list[dict], agent: object | None = None, pre_scanners: list[object] | None = None, post_scanners: list[object] | None = None) -> list[dict]
Run a guardrail red-team evaluation against a list of test cases.
Selects unit mode (scanner-only, no LLM calls) when agent is None,
and E2E mode when an agent is provided.
Args:
inputs: list of {"prompt": str, "expected_blocked": bool, "label": str}
agent: optional agent with a stream() method; omit for unit mode
pre_scanners: list of input scanner instances
post_scanners: optional list of output scanner instances (E2E only)
Returns:
list of result dicts — see :func:run_unit_eval / :func:run_e2e_eval
Example::
from latent.flows.guardrail_redteam.flow import guardrail_redteam
from latent.guardrails.scanners.builtin import LanguageScanner
results = guardrail_redteam(
inputs=[
{"prompt": "hello", "expected_blocked": False, "label": "en"},
{"prompt": "bonjour", "expected_blocked": True, "label": "fr"},
],
pre_scanners=[LanguageScanner(allowed_languages=["en"])],
)
for r in results:
print(r["label"], r["blocked"], r["score"])
run_e2e_eval¶
run_e2e_eval(test_cases: list[dict], agent: object, pre_scanners: list[object], post_scanners: list[object] | None = None) -> list[dict]
E2E mode: routes prompts through a full agent wrapped with GuardrailMiddleware.
Args: test_cases: list of {"prompt": str, "expected_blocked": bool, "label": str} agent: any agent with a stream() method pre_scanners: list of input scanner instances post_scanners: optional list of output scanner instances
Returns: list of result dicts with blocked, violated_rules, response, score fields
run_unit_eval¶
Unit mode: test each prompt against pre scanners only (no agent call).
Args: test_cases: list of {"prompt": str, "expected_blocked": bool, "label": str} pre_scanners: list of scanner instances with scan(prompt) method
Returns: list of result dicts with blocked, violated_rules, score fields