latent.rate_limiter.profiles¶
Per-provider concurrency profiles.
A profile pins the static cap (and AIMD parameters, reserved for a
future PR) for a given (provider, model) pair. Lookups use
fnmatch patterns so a single entry covers a family of models
(e.g. anthropic.claude-opus-*).
Defaults are conservative — they reflect what we've empirically seen
hold in our accounts. Override via env var or the consumer's
latent.toml when the deployment knows the actual quota.
Classes¶
ProviderProfile¶
ProviderProfile(max_concurrent: int = 50, min_concurrent: int = 1, successes_per_increment: int = 100, throttle_decay_factor: float = 0.5)
Static per-provider concurrency configuration.
Functions¶
get_profile¶
Return the profile for a given (provider, model) pair.
Lookup order: env-var overrides → built-in :data:PROFILES →
provider's default → :data:_FALLBACK. Within a provider,
fnmatch patterns are tried in declaration order — entries should
be ordered specific-first.
infer_provider¶
Best-effort provider id from a model string. Returns "unknown" if no match.
profile_key¶
Backend storage key for a (provider, model) pair.
Slot accounting is per-(provider, model) — siblings in the same
family don't share a slot pool. This is the most conservative
interpretation; if the provider's quota is actually shared across
siblings, the AIMD follow-up will converge each sibling to its fair
share independently.