Skip to content

latent.guardrails.schema_leak

Deterministic schema / SQL leakage detector.

Generic guardrail primitive for NL-to-SQL projects. Consumers supply a SchemaVocab (the tables and columns their DB exposes); the detectors here score arbitrary text against that vocab. The module has no dependency beyond re and never opens a DB, so it can be safely imported from both runtime scanners and offline eval scorers without either dragging in the other's heavy dependencies.

Two precision tiers, with strict ⊂ loose enforced by construction: detect_schema_leak (loose) runs the strict detector first and returns its match if any; only then does it apply the loose-only fallback patterns.

Classes

SchemaVocab

SchemaVocab(internal_tables: frozenset[str] = frozenset(), ambiguous_tables: frozenset[str] = frozenset(), internal_columns: frozenset[str] = frozenset(), ambiguous_columns: frozenset[str] = frozenset())

Immutable schema-identifier vocabulary consumed by both detectors.

Tables and columns are each split by the underscore heuristic: names with _ are "internal" (gate on sight, both tiers); single-word names are "ambiguous" (also read as English — gate only in SQL context, both tiers).

__post_init__ normalises every set to frozenset, validates that each element is a string, strips whitespace, and drops empty values. Then it raises if the resulting vocab is empty across all four sets. A security primitive that has nothing to look for silently detects nothing; the guard runs on every construction path (direct or from_iterables) so it cannot be bypassed with SchemaVocab(ambiguous_tables={""}) or similar whitespace-only inputs.

Functions

detect_schema_leak

detect_schema_leak(text: str, vocab: SchemaVocab) -> str | None

Loose tier — high recall. Structurally strict ⊂ loose.

Runs detect_schema_leak_strict first and returns its match on any hit — this makes the subset property a construction guarantee, not a test invariant that future strict-tier changes could accidentally break. Only when strict is clean do the loose-only fallback patterns (bare SQL-keyword IGNORECASE alternation + year-near-id proximity) run.

detect_schema_leak_strict

detect_schema_leak_strict(text: str, vocab: SchemaVocab) -> str | None

Strict tier — high precision. Gates only on high-evidence signals.

  1. Any internal (snake_case) table on sight — case-insensitive.
  2. Any internal (snake_case) column on sight — case-insensitive.
  3. Ambiguous single-word table with strong SELECT/JOIN/qualified SQL structure (see _ambiguous_table_in_sql_context).
  4. Mutating SQL (INSERT INTO / UPDATE <t> [alias] SET / DELETE FROM / TRUNCATE [TABLE] / DROP TABLE / CREATE TABLE / ALTER TABLE / MERGE INTO) followed by any target — the compound keyword+target shape is enough evidence even when the target is not in the vocabulary (catches stale or unknown tables like DROP TABLE temp_backup).
  5. SELECT … FROM combined with a structural keyword (WHERE / GROUP BY / ORDER BY). Bare "select … from" in prose does NOT gate.

Ambiguous single-word columns (title, name, price) are intentionally not gated — firing on them in SELECT price FROM … would false-positive on prose ("select the price from the menu"), and firing only inside confirmed SQL structure is redundant with rule 5.

Methods

SchemaVocab.all_columns

all_columns() -> frozenset[str]

SchemaVocab.all_tables

all_tables() -> frozenset[str]

SchemaVocab.from_iterables

from_iterables(tables: Iterable[str], columns: Iterable[str]) -> SchemaVocab