latent.guardrails.warmup¶
Functions¶
warmup¶
Force each scanner's one-time model load now, best-effort.
Call at process boot (off the request path) so the first user request does
not pay the cold-start of building the guardrail models. Scanners without a
callable preload (pure-Python inline guardrails, no-op stand-ins) are
skipped rather than reported as a warmup failure.
Never raises on a scanner's own failure: a build error is logged and the
scanner is left to build lazily on first use, so a broken model never fails
boot. No timeout is imposed, though — a hung model download hangs warmup,
so bound the call (e.g. asyncio.timeout) if boot must not stall. A
timeout stops the wait, not the build: the build keeps the process-wide
build lock until it finishes, so every other scanner's first scan waits
behind it.
(Cancellation and interpreter exit still propagate, as they must.)
The scanner cache builds one model at a time across the whole process
(transformers' loader patches process-wide state and is not thread-safe),
so scanners start together but their models load in turn, and scanners
sharing a config still load their model only once.