Skip to content

latent.rate_limiter

Provider-agnostic concurrency / rate-limiting layer for LLM calls.

Usage from agent code — drop-in for litellm.acompletion:

from latent.rate_limiter import (
    rate_limited_completion,
    rate_limited_completion_stream,
)

# Non-streaming
response = await rate_limited_completion(model="...", messages=[...])

# Streaming
async for chunk in rate_limited_completion_stream(
    model="...", messages=[...], stream=True
):
    ...

The per-provider concurrency cap comes from :mod:latent.rate_limiter.profiles; storage is the SQLite backend by default (single-host multi-process). Override the backend with LATENT_RATE_LIMIT_BACKEND; override profiles with LATENT_RATE_LIMIT_PROFILES_JSON.