Skip to content

latent.rag.adaptive

Adaptive RAG runtime patterns.

Corrective retrieval (CRAG), LLM-based query rewriting, HyDE, and confidence gating.

Classes

CachedRetriever

CachedRetriever(retriever: Any, embedding_provider: Any, similarity_threshold: float = 0.95, max_cache_size: int = 1000)

Wraps a retriever with embedding-based semantic caching.

Uses pure Python cosine similarity — no numpy dependency. Uses an OrderedDict for O(1) LRU eviction.

ConfidenceGate

ConfidenceGate(model: str = 'claude-haiku-4-5', threshold: float = 0.8, cache: bool = False)

Skip retrieval when the LLM is confident it can answer without context.

CorrectiveRetriever

CorrectiveRetriever(retriever: Any, relevance_judge: Any, fallback_retriever: Any | None = None, relevance_threshold: float = 0.5, max_corrections: int = 2, score_field: str = 'score')

Wraps a retriever with relevance evaluation and corrective actions.

Implements search() and search_with_threshold() only — not the full Retriever protocol. For indexing, use the underlying retriever.

CrossEncoderReranker

CrossEncoderReranker(model_name: str = 'cross-encoder/ms-marco-MiniLM-L-6-v2')

Reranks using a cross-encoder model from sentence-transformers. Optional [rag-reranker] extra.

HyDERewriter

HyDERewriter(model: str = 'claude-haiku-4-5', system: str | None = None, max_tokens: int | None = None)

Hypothetical Document Embeddings (HyDE) query expander.

Generates a hypothetical answer paragraph for the query and returns it alongside the original query for embedding-based retrieval.

Pass a custom system prompt to match the style, language, and domain of the target knowledge base — the closer the hypothetical document resembles real chunks, the better the embedding recall.

Set max_tokens to cap the hypothetical's length. It is only embedded, never shown, so a short cap (e.g. 96) trims generation latency with little recall impact — an uncapped answer paragraph can run several seconds.

LLMReranker

LLMReranker(model: str = 'claude-haiku-4-5')

Reranks chunks using an LLM to score (query, chunk) relevance.

QueryRewriter

QueryRewriter(model: str = 'claude-haiku-4-5')

LLM-based query rewriter that expands a query into multiple alternatives.

Reranker

Reranker()

Protocol for reranking retrieved chunks.

Functions

reorder_for_context

reorder_for_context(chunks: list[RetrievedChunk]) -> list[RetrievedChunk]

Reorder chunks so strongest are at start and end (lost-in-the-middle fix).

Example: scores [0.9, 0.8, 0.7, 0.6, 0.5] -> [0.9, 0.7, 0.5, 0.6, 0.8]

Methods

CachedRetriever.clear_cache

clear_cache() -> None

CachedRetriever.search

search(query: str, k: int = 5, kwargs = {}) -> list[RetrievedChunk]

CachedRetriever.search_with_threshold

search_with_threshold(query: str, k: int = 5, threshold: float = 0.0, kwargs = {}) -> list[RetrievedChunk]

ConfidenceGate.should_retrieve

should_retrieve(query: str) -> bool

Return True if retrieval is needed.

CorrectiveRetriever.search

search(query: str, k: int = 5) -> list[RetrievedChunk]

Retrieve, evaluate, and correct.

CorrectiveRetriever.search_with_threshold

search_with_threshold(query: str, k: int = 5, threshold: float = 0.0) -> list[RetrievedChunk]

Like search() but initial retrieval uses score threshold.

CrossEncoderReranker.rerank

rerank(query: str, chunks: list[RetrievedChunk], top_k: int | None = None) -> list[RetrievedChunk]

HyDERewriter.expand

expand(query: str) -> list[str]

Return [original_query, hypothetical_document].

HyDERewriter.expand_tokens

expand_tokens(query: str) -> list[str]

Return word tokens extracted from the hypothetical document.

LLMReranker.rerank

rerank(query: str, chunks: list[RetrievedChunk], top_k: int | None = None) -> list[RetrievedChunk]

QueryRewriter.expand

expand(query: str) -> list[str]

Return the original query plus LLM-generated rewrites.

QueryRewriter.expand_tokens

expand_tokens(query: str) -> list[str]

Return word tokens extracted from all expanded queries.

Reranker.rerank

rerank(query: str, chunks: list[RetrievedChunk], top_k: int | None = None) -> list[RetrievedChunk]