Skip to content

latent.rate_limiter.profiles

Per-provider concurrency profiles.

A profile pins the static cap (and AIMD parameters, reserved for a future PR) for a given (provider, model) pair. Lookups use fnmatch patterns so a single entry covers a family of models (e.g. anthropic.claude-opus-*).

Defaults are conservative — they reflect what we've empirically seen hold in our accounts. Override via env var or the consumer's latent.toml when the deployment knows the actual quota.

Classes

ProviderProfile

ProviderProfile(max_concurrent: int = 50, min_concurrent: int = 1, successes_per_increment: int = 100, throttle_decay_factor: float = 0.5)

Static per-provider concurrency configuration.

Functions

get_profile

get_profile(provider_id: str, model_id: str) -> ProviderProfile

Return the profile for a given (provider, model) pair.

Lookup order: env-var overrides → built-in :data:PROFILES → provider's default → :data:_FALLBACK. Within a provider, fnmatch patterns are tried in declaration order — entries should be ordered specific-first.

infer_provider

infer_provider(model: str) -> str

Best-effort provider id from a model string. Returns "unknown" if no match.

profile_key

profile_key(provider_id: str, model_id: str) -> str

Backend storage key for a (provider, model) pair.

Slot accounting is per-(provider, model) — siblings in the same family don't share a slot pool. This is the most conservative interpretation; if the provider's quota is actually shared across siblings, the AIMD follow-up will converge each sibling to its fair share independently.

Attributes

PROFILES