latent.agents.caching¶
Provider-stable prompt caching for ReActAgent.
Declare caching intent with a CachePolicy; the agent emits the correct wire markers per provider, or none, and never crashes on an unknown one.
Classes¶
CachePolicy¶
What to cache. Pass to the agent to enable (default: off).
SystemContent¶
System message split into a cacheable prefix and an uncached tail.
Only static ever carries the cache marker, so per-turn content in dynamic can never leak into the cached prefix.
Functions¶
cache_tool_defs¶
cache_tool_defs(tool_defs: list[dict[str, Any]], policy: CachePolicy | None, model: str) -> list[dict[str, Any]]
Attach a cache marker to the tool array when the policy enables it.
emits_cache_markers¶
Whether the provider for model accepts inline cache markers.
mark_message_cache_breakpoint¶
mark_message_cache_breakpoint(llm_messages: list[dict[str, Any]], policy: CachePolicy | None, model: str) -> None
Attach a rolling cache marker to the last message, in place.
ReAct re-sends the full, growing message list on every step, so without a breakpoint the accumulated tool results are re-tokenized at full price each step — cost is quadratic in step count. Marking the last message makes each step write a cache entry for the new prefix that the next step reads at the cache-read discount. Reads auto-match the longest previously-written prefix, so only the newest write point needs a marker; we strip any prior message marker first to stay within the provider's breakpoint budget (system + tools + one rolling message = 3 of 4).
render_system_message¶
render_system_message(content: str | SystemContent | list[dict[str, Any]], policy: CachePolicy | None, model: str) -> dict[str, Any] | None
Build the system message dict, attaching cache markers per policy.
Returns None when there is no system content to send.