Skip to content

latent.guardrails.scanners.context

Context engineering scanners — validators and compactors for LLM message arrays.

Each scanner implements the MessageInputScanner duck-type (name, on_error, and scan_messages(messages, tools)).

Static validators are fast (<5 ms, no LLM / network calls). Active compactors may return rewritten_messages to transform context.

Classes

CompactionScanner

CompactionScanner(trigger_utilization: float = 0.8, target_utilization: float = 0.6, model_limit: int = 200000, on_error: OnError = 'ignore')

Triggers context compaction when token budget exceeds threshold.

CompressionQualityAuditor

CompressionQualityAuditor(model: str = 'gpt-4o-mini', probes: list[str] | None = None, on_error: OnError = 'ignore')

Evaluates quality of compressed context using probe questions.

ContradictionDetector

ContradictionDetector(model: str = 'gpt-4o-mini', on_error: OnError = 'ignore')

Detects conflicting information across context sources.

DistractionScorer

DistractionScorer(model: str = 'gpt-4o-mini', on_error: OnError = 'ignore')

Scores context elements for relevance to the current task.

HistoryBloatDetector

HistoryBloatDetector(max_history_pct: float = 0.6, on_error: OnError = 'ignore')

Fails if user/assistant history exceeds max_history_pct of total tokens.

Uses :func:budget_breakdown for categorisation.

KVCacheStabilityAuditor

KVCacheStabilityAuditor(on_error: OnError = 'ignore')

Checks system prompt for cache-breaking dynamic elements.

Detects ISO timestamps, UUIDs, session IDs, request counters, and version numbers with patch components (e.g. v1.2.3) via regex.

  • passed = no cache-breakers found
  • metadata["cache_breakers"] = list of {pattern, matched_text}

MiddleContentDetector

MiddleContentDetector(critical_patterns: list[str] | None = None, min_total_tokens: int = 4000, on_error: OnError = 'ignore')

Flags critical instructions in the attention trough (10-90th percentile).

Concatenates all message content, locates critical_patterns via case-insensitive word-boundary search, and checks whether any fall in the middle 80% of the total token count.

  • score = fraction of critical patterns found in the middle
  • passed = no critical patterns in the middle
  • metadata["findings"] = list of {pattern, position_pct, message_role}

ObservationMaskingScanner

ObservationMaskingScanner(keep_last_n: int = 3, on_error: OnError = 'ignore')

Compresses consumed tool outputs, keeping recent ones intact.

PoisoningDetector

PoisoningDetector(model: str = 'gpt-4o-mini', on_error: OnError = 'ignore')

Detects hallucinated facts re-entering context.

SummaryInjectionScanner

SummaryInjectionScanner(every_n_messages: int = 20, on_error: OnError = 'ignore')

Injects periodic summary messages into long conversations.

SystemPromptStructureAuditor

SystemPromptStructureAuditor(on_error: OnError = 'ignore')

Checks system prompt for structural best practices.

Finds the first system message and checks:

  • Has identity/role statement in first 200 chars
  • Critical constraints appear in first 10% AND last 10% (edge-anchoring)
  • No mixing of high-level directives with implementation details (altitude check)

  • passed = no structural issues

  • metadata["issues"] = list of issue descriptions

TokenBudgetAuditor

TokenBudgetAuditor(model_limit: int = 200000, warn_pct: float = 0.7, compact_pct: float = 0.8, critical_pct: float = 0.9, on_error: OnError = 'ignore')

Checks overall context utilisation against model limits.

Uses :func:budget_breakdown to compute utilisation.

  • score = utilisation ratio (0.0 -- 1.0)
  • passed = utilisation < compact_pct
  • metadata includes full :class:BudgetBreakdown fields + severity level

ToolDescriptionLinter

ToolDescriptionLinter(on_error: OnError = 'ignore')

Validates tool definitions for quality.

Checks each tool for:

  • Has description (>10 chars)
  • Description mentions what it returns / what output to expect
  • Parameter count <= 8 (flags over-consolidated tools)
  • Naming follows verb_noun pattern (optional, metadata only)

If tools is None or empty, passes immediately.

  • score = fraction of tools with issues
  • metadata["tool_findings"] = list of {tool_name, issues: list[str]}

ToolOutputOffloadScanner

ToolOutputOffloadScanner(max_output_tokens: int = 2000, scratch_dir: str | None = None, on_error: OnError = 'ignore')

Offloads large tool outputs to scratch files.

Methods

CompactionScanner.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

CompressionQualityAuditor.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict] | None = None) -> ScanResult

ContradictionDetector.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict] | None = None) -> ScanResult

DistractionScorer.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict] | None = None) -> ScanResult

HistoryBloatDetector.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

KVCacheStabilityAuditor.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

MiddleContentDetector.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

ObservationMaskingScanner.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

PoisoningDetector.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict] | None = None) -> ScanResult

SummaryInjectionScanner.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

SystemPromptStructureAuditor.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

TokenBudgetAuditor.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

ToolDescriptionLinter.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

ToolOutputOffloadScanner.scan_messages

scan_messages(messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None = None) -> ScanResult

ToolOutputOffloadScanner.scratch_dir

Lazily create the scratch directory on first access.