latent.rate_limiter.protocol¶
Backend protocol for the provider-agnostic rate limiter.
A backend stores per-key concurrency counters. Implementations may keep
state in-process (a single asyncio.Semaphore), in a single-host
multi-process store (SQLite + WAL, the default), or in a distributed
store (Redis). The wrapper in :mod:latent.rate_limiter.wrapper does
not know which backend it's using.
Keys are opaque strings (typically "{provider}:{model}"); capacities
are positive integers. Acquire blocks until a slot is available;
release returns a slot. report_throttle is a one-way signal sent
when the underlying provider returns litellm.RateLimitError — a
backend may use it for adaptive cap adjustment (AIMD) in a future PR;
the current PR's backend just records the count for observability.
Classes¶
RateLimiterBackend¶
Storage + coordination primitive for per-key concurrency caps.
Methods¶
RateLimiterBackend.acquire¶
Block until a slot for key is available, then take it.
max_concurrent is the cap used on first-seen keys. Subsequent
calls with the same key ignore the value — the cap is fixed at
first observation. Use :meth:set_capacity to change it later.
RateLimiterBackend.current_capacity¶
Return the cap currently in force for key (0 if unknown).
RateLimiterBackend.list_keys¶
Return all keys the backend currently knows about.
Used by the MLflow metric emitter at end-of-flow to log per-key throttle counts without the caller needing to know which keys were touched.
RateLimiterBackend.release¶
Return a previously-acquired slot for key. Idempotent on extra calls.
RateLimiterBackend.report_throttle¶
Record that litellm.RateLimitError was observed for key.
Backends may use this for adaptive cap adjustment; the static backend just increments a counter for observability.
RateLimiterBackend.throttle_count¶
Cumulative count of report_throttle calls for key.