Agentic retrieval transforms how enterprise AI systems handle complex queries by breaking them into sub-queries, ensuring more accurate and contextually rich responses. Unlike standard retrieval, which returns concise but potentially incomplete answers, agentic retrieval plans and executes multiple searches to gather sufficient evidence. This approach is particularly valuable for multi-part questions, where a single query vector might not capture all user intents.
How Agentic Retrieval Works
Agentic retrieval operates within a planning loop, where the system dynamically generates sub-queries, assesses the evidence, and decides whether additional searches are needed. This process ensures that the final response is both comprehensive and relevant. For example, when a user asks a support assistant to compare two products across three dimensions, agentic retrieval treats this as six simultaneous questions. It then retrieves and synthesizes information for each dimension, providing a detailed and accurate answer.
Amazon Bedrock Managed Knowledge Base integrates this capability through its AgenticRetrieveStream API, which streams trace events back to the application. These events show the model's retrieval plan, allowing developers to understand and optimize the search process. The API is part of Amazon Bedrock's fully managed RAG (Retrieval Augmented Generation) capability, which handles chunking, embedding, storage, and retrieval without requiring a self-managed vector store.
Implementation with LangChain
To implement agentic retrieval, developers can use the langchain-aws package, which exposes both standard and agentic retrieval methods. The package integrates seamlessly with Amazon Bedrock Managed Knowledge Base, enabling applications to leverage its planning loop. For instance, the AgenticRetrieveStream API can be used to handle multi-part questions, while the Retrieve API serves single-shot queries.
Here’s a code snippet demonstrating how to create a knowledge base and attach an S3 bucket as a data source:
```python import boto3 import os REGION = os.environ["AWS_REGION"] bedrock_agent = boto3.client("bedrock-agent", region_name=REGION) response = bedrock_agent.create_knowledge_base( name=KB_NAME, roleArn=KB_ROLE_ARN, knowledgeBaseConfiguration={ "type": "MANAGED", "managedKnowledgeBaseConfiguration": { "embeddingModelType": "MANAGED", }, }, ) KB_ID = response["knowledgeBase"]["knowledgeBaseId"] ```
After creating the knowledge base, an ingestion job is started to populate it with data. The wait_for_ingestion function polls the job status until completion, ensuring data is ready for retrieval.
Cost Efficiency and Use Cases
While agentic retrieval provides more thorough responses, it incurs higher costs due to multiple searches and foundation model (FM) inference. However, it is the right choice for complex queries where completeness is critical. For simpler queries, standard retrieval offers a cost-effective alternative.
Practical use cases include enterprise support assistants, where users often ask multi-part questions requiring detailed answers. By combining agentic retrieval with custom evaluators, businesses can ensure responses are not only linguistically accurate but also aligned with domain-specific rules and constraints.
Permissions and Security
Implementing agentic retrieval requires careful permission management. The knowledge base assumes a service role to access documents and call embedding models, while the application uses an AWS STS caller identity to query. The service role needs permissions like s3:ListBucket and s3:GetObject, scoped to the resource account. The caller identity requires bedrock:AgenticRetrieveStream and bedrock:Retrieve permissions, scoped to the knowledge base ARN.
For example, the service role trust policy might look like this:
```json { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Principal": {"Service": "bedrock.amazonaws.com"}, "Action": "sts:AssumeRole", "Condition": { "StringEquals": {"aws:SourceAccount": "111122223333"}, "ArnLike": { "aws:SourceArn": "arn:aws:bedrock:us-east-1:111122223333:knowledge-base/*" } } }] } ```
Conclusion
Agentic retrieval is a powerful tool for enhancing enterprise AI applications, particularly when handling complex queries. By integrating Amazon Bedrock Managed Knowledge Base with LangChain, developers can build systems that provide accurate, contextually rich responses. While it comes with higher costs, its ability to ensure completeness and relevance makes it invaluable for critical enterprise use cases.