latent.rag.adaptive¶
Adaptive RAG runtime patterns.
Corrective retrieval (CRAG), LLM-based query rewriting, HyDE, and confidence gating.
Classes¶
CachedRetriever¶
CachedRetriever(retriever: Any, embedding_provider: Any, similarity_threshold: float = 0.95, max_cache_size: int = 1000)
Wraps a retriever with embedding-based semantic caching.
Uses pure Python cosine similarity — no numpy dependency. Uses an OrderedDict for O(1) LRU eviction.
ConfidenceGate¶
Skip retrieval when the LLM is confident it can answer without context.
CorrectiveRetriever¶
CorrectiveRetriever(retriever: Any, relevance_judge: Any, fallback_retriever: Any | None = None, relevance_threshold: float = 0.5, max_corrections: int = 2, score_field: str = 'score')
Wraps a retriever with relevance evaluation and corrective actions.
Implements search() and search_with_threshold() only — not the full Retriever protocol. For indexing, use the underlying retriever.
CrossEncoderReranker¶
Reranks using a cross-encoder model from sentence-transformers. Optional [rag-reranker] extra.
HyDERewriter¶
HyDERewriter(model: str = 'claude-haiku-4-5', system: str | None = None, max_tokens: int | None = None)
Hypothetical Document Embeddings (HyDE) query expander.
Generates a hypothetical answer paragraph for the query and returns it alongside the original query for embedding-based retrieval.
Pass a custom system prompt to match the style, language, and domain
of the target knowledge base — the closer the hypothetical document
resembles real chunks, the better the embedding recall.
Set max_tokens to cap the hypothetical's length. It is only embedded,
never shown, so a short cap (e.g. 96) trims generation latency with little
recall impact — an uncapped answer paragraph can run several seconds.
LLMReranker¶
Reranks chunks using an LLM to score (query, chunk) relevance.
QueryRewriter¶
LLM-based query rewriter that expands a query into multiple alternatives.
Reranker¶
Protocol for reranking retrieved chunks.
Functions¶
reorder_for_context¶
Reorder chunks so strongest are at start and end (lost-in-the-middle fix).
Example: scores [0.9, 0.8, 0.7, 0.6, 0.5] -> [0.9, 0.7, 0.5, 0.6, 0.8]
Methods¶
CachedRetriever.clear_cache¶
CachedRetriever.search¶
CachedRetriever.search_with_threshold¶
search_with_threshold(query: str, k: int = 5, threshold: float = 0.0, kwargs = {}) -> list[RetrievedChunk]
ConfidenceGate.should_retrieve¶
Return True if retrieval is needed.
CorrectiveRetriever.search¶
Retrieve, evaluate, and correct.
CorrectiveRetriever.search_with_threshold¶
Like search() but initial retrieval uses score threshold.
CrossEncoderReranker.rerank¶
HyDERewriter.expand¶
Return [original_query, hypothetical_document].
HyDERewriter.expand_tokens¶
Return word tokens extracted from the hypothetical document.
LLMReranker.rerank¶
QueryRewriter.expand¶
Return the original query plus LLM-generated rewrites.
QueryRewriter.expand_tokens¶
Return word tokens extracted from all expanded queries.