Skip to content

latent.rag.chroma

ChromaDB-backed vector retriever.

Classes

ChromaRetriever

ChromaRetriever(collection_name: str = 'knowledge', persist_directory: str = './chroma_db', embedding_provider: EmbeddingProvider | str | None = None, embedding_model: str | None = None, distance_fn: str = 'cosine', query_expander: QueryExpander | None = None, high_relevance_threshold: float | None = None, medium_relevance_threshold: float | None = None, client: Any | None = None)

Retriever backed by ChromaDB with persistent storage.

Parameters

collection_name: Name of the ChromaDB collection. persist_directory: Path to ChromaDB's persistent storage. embedding_provider: An EmbeddingProvider instance or a provider name string (resolved lazily via :func:create_embedding_provider). embedding_model: Model name — only used when embedding_provider is a string. distance_fn: Distance metric ("cosine", "l2", "ip"). Note: search_with_threshold scores are only meaningful with "cosine" (similarity in [0, 1]). query_expander: Optional QueryExpander for expanding queries before search.

Functions

deduplicate_chunks

deduplicate_chunks(chunks: list[RetrievedChunk], k: int) -> list[RetrievedChunk]

Deduplicate chunks by ID, keeping the highest score, return top-k.

Methods

ChromaRetriever.add_chunks

add_chunks(chunks: list[DocumentChunk]) -> None

Add pre-built DocumentChunk objects.

ChromaRetriever.add_texts

add_texts(texts: list[str], metadatas: list[dict[str, Any]] | None = None) -> None

Add texts to the collection with content-based IDs.

ChromaRetriever.count

count() -> int

Return number of documents in the collection.

ChromaRetriever.delete_collection

delete_collection() -> None

Delete and re-create the collection (clears all data).

ChromaRetriever.get_all_chunks

get_all_chunks() -> list[DocumentChunk]

Return all indexed chunks from the collection.

ChromaRetriever.search

search(query: str, k: int = 5) -> list[RetrievedChunk]

Semantic search returning top-k chunks.

ChromaRetriever.search_multi

search_multi(queries: list[str], k: int = 5) -> list[RetrievedChunk]

Search with multiple queries, deduplicate by ID, return top-k.

ChromaRetriever.search_with_threshold

search_with_threshold(query: str, k: int = 5, threshold: float = 0.0) -> list[RetrievedChunk]

Semantic search filtered by threshold.