latent.rate_limiter.metrics¶
MLflow integration for rate-limiter observability.
Emits per-(provider, model) throttle counts as MLflow metrics so
back-pressure becomes visible per eval run without operators having to
grep logs or query CloudWatch by hand.
Called from :func:latent.flows.agent_inference_flow.agent_inference_flow
at end-of-run. No-op when MLflow isn't installed or no active run.
Functions¶
log_rate_limiter_metrics_to_mlflow¶
Emit per-key throttle counts to the active MLflow run, if any.
Safe to call when MLflow isn't installed or no run is active — both cases short-circuit silently. Errors during emission are logged but not raised; observability shouldn't break the flow.
Metric names emitted:
- rate_limiter.{sanitized_key}.throttle_count
- rate_limiter.{sanitized_key}.capacity