Skip to content

latent.context.tokens

Token estimation utilities.

Uses UTF-8 byte count // 4 as a multilingual-safe approximation. This matches the approach in latent.guardrails.scanners.builtin.TokenLimitScanner.

Functions

estimate_tokens

estimate_tokens(text: str, method: Literal['utf8'] = 'utf8') -> int

Estimate token count from text.

The "utf8" method divides the UTF-8 byte length by 4, which is a reasonable approximation across Latin and non-Latin scripts (e.g. Hebrew characters are 2 bytes each, so a pure char-count would under-estimate by ~2x).

Parameters

text: The input string to estimate tokens for. method: Estimation strategy. Currently only "utf8" is supported.

Returns

int Estimated token count (always >= 0).

message_text

message_text(msg: dict[str, Any]) -> str

Extract the textual content of a message for token estimation.

Handles both plain string content and the OpenAI multimodal list-of-content-parts format.

Parameters

msg: An OpenAI-style message dict (must have "content").

Returns

str The extracted text content.