Skip to content

latent.rag.bm25

BM25 keyword retriever using rank-bm25.

Classes

BM25Retriever

BM25Retriever(k1: float = 1.5, b: float = 0.75, tokenizer: Tokenizer | None = None, bm25_tokenizer: str | None = None)

BM25Okapi keyword retriever.

Parameters

k1: BM25 term-frequency saturation parameter. b: BM25 document-length normalisation parameter. tokenizer: Optional custom tokenizer (str) -> list[str]. When None the default whitespace/punctuation splitter is used. Pass a language-aware tokenizer (e.g. Hebrew prefix stripper) to improve recall for morphologically rich languages. bm25_tokenizer: Named tokenizer from TOKENIZER_REGISTRY (e.g. "default", "hebrew"). Takes precedence over tokenizer when set.

Methods

BM25Retriever.add_chunks

add_chunks(chunks: list[DocumentChunk]) -> None

Add chunks to the existing index (rebuilds with accumulated docs).

BM25Retriever.add_texts

add_texts(texts: list[str], metadatas: list[dict] | None = None) -> None

Index plain texts (builds content-based IDs).

BM25Retriever.delete_collection

delete_collection() -> None

Clear the index.

BM25Retriever.index

index(chunks: list[DocumentChunk]) -> None

Build (or rebuild) the BM25 index from chunks.

BM25Retriever.indexed

Whether the index has been built.

BM25Retriever.search

search(query: str, k: int = 5) -> list[RetrievedChunk]

Return top-k chunks by BM25 score.

BM25Retriever.search_multi

search_multi(queries: list[str], k: int = 5) -> list[RetrievedChunk]

Search with multiple queries, deduplicate by ID, return top-k.

BM25Retriever.search_with_threshold

search_with_threshold(query: str, k: int = 5, threshold: float = 0.0) -> list[RetrievedChunk]

Return top-k chunks with score >= threshold.

Attributes

TOKENIZER_REGISTRY

Tokenizer