Skip to content

latent.rag.hybrid

Hybrid retriever combining semantic (ChromaDB) and keyword (BM25) search via RRF.

Classes

HybridRetriever

HybridRetriever(chroma: ChromaRetriever, bm25: BM25Retriever | None = None, alpha: float = 0.5, rrf_k: int = 60, query_expander: QueryExpander | None = None, semantic_expander: QueryExpander | None = None, tokenizer = None, bm25_tokenizer: str | None = None)

Reciprocal Rank Fusion of a :class:ChromaRetriever and a :class:BM25Retriever.

Parameters

chroma: Semantic vector retriever. bm25: Keyword retriever. If None a fresh instance is created (and lazily indexed from the Chroma collection on first search). alpha: Weight of the semantic score in [0, 1]. 1 - alpha weights the keyword score. rrf_k: RRF smoothing constant (default 60). query_expander: Optional query expander applied to BM25 queries. semantic_expander: Optional query expander applied to the semantic (embedding) side. Use :class:~latent.rag.adaptive.HyDERewriter here to generate a hypothetical answer document that is embedded alongside the original query — improving recall for questions whose surface form differs from the answer's phrasing in the knowledge base. tokenizer: Optional custom tokenizer for BM25. Forwarded to the BM25Retriever when bm25 is None. Use a language-aware tokenizer (e.g. Hebrew prefix normalisation) to improve keyword recall. bm25_tokenizer: Named tokenizer from TOKENIZER_REGISTRY. Forwarded to the BM25Retriever when bm25 is None.

Methods

HybridRetriever.add_chunks

add_chunks(chunks: list[DocumentChunk]) -> None

Add document chunks to both indexes.

HybridRetriever.add_texts

add_texts(texts: list[str], metadatas: list[dict] | None = None) -> None

Add texts to both Chroma and BM25 indexes.

HybridRetriever.delete_collection

delete_collection() -> None

Delete the Chroma collection and clear BM25 index.

HybridRetriever.search

search(query: str, k: int = 5) -> list[RetrievedChunk]

Fused semantic + keyword search returning top-k.

Thin non-streaming wrapper over :meth:search_events: drains the phase observability events and returns only the fused result. The per-phase event objects are cheap next to the embedding / LLM work they bracket.

HybridRetriever.search_events

search_events(query: str, k: int = 5) -> AsyncGenerator[Any, None]

Streaming variant of :meth:search for phase-level observability.

Yields a nested RetrieverStart / RetrieverEnd pair around each retrieval phase — hyde (semantic query expansion, usually an LLM call), dense (embedding + vector search), sparse (BM25) — then yields the fused list[RetrievedChunk] as the single non-event item. Contract mirrors @retriever; that decorator stamps parent_run_id so the phase spans nest under the enclosing retriever span.

A phase whose backend await raises still emits its RetrieverEnd (empty documents) before the exception propagates, so no phase span is left open — matching the (start, end)-pair invariant @retriever guarantees for the enclosing span.

HybridRetriever.search_multi

search_multi(queries: list[str], k: int = 5) -> list[RetrievedChunk]

Search with multiple queries, boost multi-query hits, return top-k.

Chunks found by multiple queries are stronger relevance signals — they match the user's intent from different phrasings. We apply a multiplicative boost (_MULTI_QUERY_BOOST per additional query) so these chunks are more likely to cross the HIGH-relevance threshold and get their constraints/facts extracted.

Thin non-streaming drain of :meth:search_multi_events.

HybridRetriever.search_multi_events

search_multi_events(queries: list[str], k: int = 5) -> AsyncGenerator[Any, None]

Streaming, phase-emitting variant of :meth:search_multi.

Runs each query's :meth:search_events concurrently (as search_multi does) and merges their phase events (hyde/dense/sparse) into one stream via a fan-in queue, so the phases of all query variations interleave as they occur — each nesting under the enclosing retriever span. Yields the boosted + deduped fused list[RetrievedChunk] as the single non-event item, last.

HybridRetriever.search_with_threshold

search_with_threshold(query: str, k: int = 5, threshold: float = 0.0) -> list[RetrievedChunk]

Fused search filtered by score threshold.