Skip to content

latent.agents.rag_agent

RAGAgent — ReActAgent with config-driven RAG pipeline.

Downstream agents subclass RAGAgent and provide rag_config instead of implementing reranker/post-retrieval wiring logic themselves.

Classes

RAGAgent

RAGAgent(name: str, model: str, rag_config: dict[str, Any] | RAGAgentConfig, system_prompt: str | None = None, tools: list[Callable] | None = None, temperature: float = 0.0, max_tokens: int = 4096, max_iterations: int = 10, response_format: type[BaseModel] | dict | None = None, optimized_prompt: str | Path | None = None, kwargs: Any = {})

ReActAgent with an integrated RAG pipeline.

Accepts a rag_config (dict or RAGAgentConfig) that declaratively configures chunking, embeddings, backend, reranker, post-retrieval strategy, query expansion, confidence gating, and caching.

The pipeline is built lazily on first index_documents() call. A search_knowledge_base tool is auto-registered so the LLM can retrieve context during conversation.

Example::

agent = RAGAgent(
    name="support",
    model="gpt-4o",
    rag_config={
        "embeddings": {"provider": "voyage", "model": "voyage-3"},
        "reranker": {"type": "cross_encoder"},
        "post_retrieval": {"type": "reorder"},
    },
)
await agent.index_documents(docs)
response = await agent.run([Message(role="user", content="How do I reset?")])

Methods

RAGAgent.index_documents

index_documents(texts: list[str], sources: list[str] | None = None) -> int

Chunk, embed, and index documents into the RAG pipeline.

Builds the pipeline on first call using rag_config.

RAGAgent.on_session_start

on_session_start() -> dict[str, Any]

Include RAG config in session metadata.

RAGAgent.reset

reset() -> None

Reset agent state. Does NOT clear the indexed documents.

RAGAgent.search

search(query: str, k: int | None = None) -> list[RetrievedChunk]

Search the indexed documents.

Raises RuntimeError if called before index_documents().