Skip to content

latent.rag.adapter

Config-driven RAG entry point.

Classes

RAGAdapter

RAGAdapter(backend: str = 'hybrid', collection_name: str = 'knowledge', persist_directory: str = './chroma_db', embedding_provider: EmbeddingProvider | str | None = None, embedding_model: str | None = None, distance_fn: str = 'cosine', alpha: float = 0.5, rrf_k: int = 60, top_k: int = 5, score_threshold: float = 0.0, chunk_max_size: int = 1000, chunk_overlap: int = 100, query_expander: QueryExpander | None = None, cache: bool = False, cache_threshold: float = 0.95, bm25_tokenizer: str | None = None)

High-level, config-driven RAG adapter.

Instantiate directly or use :meth:from_params with a flat dict (e.g. loaded from parameters.yaml).

Parameters

backend: "chroma", "hybrid", or "bm25". collection_name: ChromaDB collection name (chroma/hybrid only). persist_directory: ChromaDB persistence path (chroma/hybrid only). embedding_provider: Provider name or EmbeddingProvider instance. Required for "chroma" and "hybrid" backends. embedding_model: Model identifier forwarded to the embedding provider. distance_fn: ChromaDB distance metric (default "cosine"). alpha: Semantic weight in [0, 1] for hybrid RRF (default 0.5). rrf_k: RRF smoothing constant (default 60). top_k: Default number of results to return. score_threshold: Default minimum score for search_with_threshold. chunk_max_size: Maximum characters per chunk. chunk_overlap: Overlap characters between chunks. query_expander: Optional QueryExpander instance for expanding search queries. bm25_tokenizer: Named tokenizer from TOKENIZER_REGISTRY (e.g. "hebrew"). Forwarded to BM25-based retrievers.

Methods

RAGAdapter.count

count() -> int

Return number of indexed documents/chunks.

RAGAdapter.from_params

from_params(params_dict: dict[str, Any]) -> RAGAdapter

Create a RAGAdapter from a flat parameter dict.

The dict keys map 1:1 to constructor arguments. Unknown keys are silently ignored so you can pass an entire parameters.yaml section.

RAGAdapter.index_chunks

index_chunks(chunks: list[DocumentChunk]) -> int

Index pre-built DocumentChunk objects.

Deduplicates by chunk ID before indexing (content-hash IDs can collide when the same text appears in multiple documents). Returns the number of unique chunks indexed.

RAGAdapter.index_documents

index_documents(texts: list[str], sources: list[str] | None = None) -> int

Chunk and index raw markdown documents.

Returns the number of chunks indexed.

RAGAdapter.search

search(query: str, top_k: int | None = None) -> list[RetrievedChunk]

Search across the configured backend.

RAGAdapter.search_multi

search_multi(queries: list[str], top_k: int | None = None) -> list[RetrievedChunk]

Search with multiple queries, deduplicate, return top results.

RAGAdapter.search_with_threshold

search_with_threshold(query: str, top_k: int | None = None, threshold: float | None = None) -> list[RetrievedChunk]

Search with a minimum score threshold.