latent.rag.chroma¶
ChromaDB-backed vector retriever.
Classes¶
ChromaRetriever¶
ChromaRetriever(collection_name: str = 'knowledge', persist_directory: str = './chroma_db', embedding_provider: EmbeddingProvider | str | None = None, embedding_model: str | None = None, distance_fn: str = 'cosine', query_expander: QueryExpander | None = None, high_relevance_threshold: float | None = None, medium_relevance_threshold: float | None = None, client: Any | None = None)
Retriever backed by ChromaDB with persistent storage.
Parameters¶
collection_name:
Name of the ChromaDB collection.
persist_directory:
Path to ChromaDB's persistent storage.
embedding_provider:
An EmbeddingProvider instance or a provider name string
(resolved lazily via :func:create_embedding_provider).
embedding_model:
Model name — only used when embedding_provider is a string.
distance_fn:
Distance metric ("cosine", "l2", "ip").
Note: search_with_threshold scores are only meaningful
with "cosine" (similarity in [0, 1]).
query_expander:
Optional QueryExpander for expanding queries before search.
Functions¶
deduplicate_chunks¶
Deduplicate chunks by ID, keeping the highest score, return top-k.
Methods¶
ChromaRetriever.add_chunks¶
Add pre-built DocumentChunk objects.
ChromaRetriever.add_texts¶
Add texts to the collection with content-based IDs.
ChromaRetriever.count¶
Return number of documents in the collection.
ChromaRetriever.delete_collection¶
Delete and re-create the collection (clears all data).
ChromaRetriever.get_all_chunks¶
Return all indexed chunks from the collection.
ChromaRetriever.search¶
Semantic search returning top-k chunks.
ChromaRetriever.search_multi¶
Search with multiple queries, deduplicate by ID, return top-k.
ChromaRetriever.search_with_threshold¶
Semantic search filtered by threshold.