latent.rate_limiter¶
Provider-agnostic concurrency / rate-limiting layer for LLM calls.
Usage from agent code — drop-in for litellm.acompletion:
from latent.rate_limiter import (
rate_limited_completion,
rate_limited_completion_stream,
)
# Non-streaming
response = await rate_limited_completion(model="...", messages=[...])
# Streaming
async for chunk in rate_limited_completion_stream(
model="...", messages=[...], stream=True
):
...
The per-provider concurrency cap comes from
:mod:latent.rate_limiter.profiles; storage is the SQLite backend
by default (single-host multi-process). Override the backend with
LATENT_RATE_LIMIT_BACKEND; override profiles with
LATENT_RATE_LIMIT_PROFILES_JSON.