Skip to content

latent.rate_limiter.metrics

MLflow integration for rate-limiter observability.

Emits per-(provider, model) throttle counts as MLflow metrics so back-pressure becomes visible per eval run without operators having to grep logs or query CloudWatch by hand.

Called from :func:latent.flows.agent_inference_flow.agent_inference_flow at end-of-run. No-op when MLflow isn't installed or no active run.

Functions

log_rate_limiter_metrics_to_mlflow

log_rate_limiter_metrics_to_mlflow(backend: RateLimiterBackend | None = None) -> None

Emit per-key throttle counts to the active MLflow run, if any.

Safe to call when MLflow isn't installed or no run is active — both cases short-circuit silently. Errors during emission are logged but not raised; observability shouldn't break the flow.

Metric names emitted: - rate_limiter.{sanitized_key}.throttle_count - rate_limiter.{sanitized_key}.capacity