An advanced Retrieval-Augmented Generation (RAG) solution designed to tackle complex questions that simple semantic similarity-based retrieval cannot solve. This project showcases a sophisticated deterministic graph acting as the "brain" of a highly controllable autonomous agent capable of answering non-trivial questions from your own data.
The full reference: a 400-page visual guide that goes deeper than any notebook can. The intuition behind every technique, side-by-side comparisons of when each one wins (and when it quietly fails), and diagrams that make the tricky parts finally click.
1,500+ copies sold · Hit #1 in Generative AI on Amazon at launch · ⭐ 4.6 stars
📖 PDF + EPUB · GitHub community price: 33% off with code RAGKING
Prompt to Production - my full course on building software with AI the way professionals do: the methods and paradigms behind reliable, efficient, modular production systems, taught systematically. 17 modules, each pairing a video lecture with a hands-on lab, from your first structured prompt to a working production system.
The course is live. Every module is out, lecture and lab.
| 🎬 7-minute video lecture |
🛠️ Hands-on tutorial |
🤖 AI assistant inside Claude Code |
One npm install adds the module's AI assistant to your Claude Code, and it guides you through the tutorial as you build.
🚀 Level up with my Agents Towards Production repository. It delivers horizontal, code-first tutorials that cover every tool and step in the lifecycle of building production-grade GenAI agents, guiding you from spark to scale with proven patterns and reusable blueprints for real-world launches, making it the smartest place to start if you're serious about shipping agents to production.
📚 Explore my comprehensive guide on RAG techniques to complement this advanced agent implementation with many other RAG techniques.
🤖 Explore my GenAI Agents Repository to complement this advanced agent implementation with many other AI Agents implementations and tutorials.
| 🚀 Cutting-edge Updates |
💡 Expert Insights |
🎯 Top 0.1% Content |
Join over 20,000 of AI enthusiasts getting unique cutting-edge insights and free tutorials! Plus, subscribers get exclusive early access and special 33% discounts to my book and the upcoming RAG Techniques course!
- Sophisticated Deterministic Graph: Acts as the "brain" of the agent, enabling complex reasoning.
- Controllable Autonomous Agent: Capable of answering non-trivial questions from custom datasets.
- Hallucination Prevention: Ensures answers are solely based on provided data, avoiding AI hallucinations.
- Multi-step Reasoning: Breaks down complex queries into manageable sub-tasks.
- Adaptive Planning: Continuously updates its plan based on new information.
- Performance Evaluation: Utilizes
Ragasmetrics for comprehensive quality assessment.
- PDF Loading and Processing: Load PDF documents and split them into chapters.
- Text Preprocessing: Clean and preprocess the text for better summarization and encoding.
- Summarization: Generate extensive summaries of each chapter using large language models.
- Book Quotes Database Creation: Create a database for specific questions that will need access to quotes from the book.
- Vector Store Encoding: Encode the book content and chapter summaries into vector stores for efficient retrieval.
- Question Processing:
- Anonymize the question by replacing named entities with variables.
- Generate a high-level plan to answer the anonymized question.
- De-anonymize the plan and break it down into retrievable or answerable tasks.
- Task Execution:
- For each task, decide whether to retrieve information or answer based on context.
- If retrieving, fetch relevant information from vector stores and distill it.
- If answering, generate a response using chain-of-thought reasoning.
- Verification and Re-planning:
- Verify that generated content is grounded in the original context.
- Re-plan remaining steps based on new information.
- Final Answer Generation: Produce the final answer using accumulated context and chain-of-thought reasoning.
The solution is evaluated using Ragas metrics:
- Answer Correctness
- Faithfulness
- Answer Relevancy
- Context Recall
- Answer Similarity
The algorithm was tested using the first Harry Potter book, allowing for monitoring of the model's reliance on retrieved information versus pre-trained knowledge. This choice enables us to verify whether the model is using its pre-trained knowledge or strictly relying on the retrieved information from vector stores.
Q: How did the protagonist defeat the villain's assistant?
To solve this question, the following steps are necessary:
- Identify the protagonist of the plot.
- Identify the villain.
- Identify the villain's assistant.
- Search for confrontations or interactions between the protagonist and the villain.
- Deduce the reason that led the protagonist to defeat the assistant.
The agent's ability to break down and solve such complex queries demonstrates its sophisticated reasoning capabilities.
- Python 3.8+
- API key for your chosen LLM provider
- Clone the repository:
git clone https://github.com/NirDiamant/Controllable-RAG-Agent.git cd Controllable-RAG-Agent - Set up environment variables:
Create a
.envfile in the root directory with your API key:you can look at theOPENAI_API_KEY= GROQ_API_KEY=.env.examplefile for reference.
- run the following command to build the docker image
docker-compose up --build
- Install required packages:
pip install -r requirements.txt
-
Explore the step-by-step tutorial:
sophisticated_rag_agent_harry_potter.ipynb -
Run real-time agent visualization (no docker):
streamlit run simulate_agent.py
-
Run real-time agent visualization (with docker): open your browser and go to
http://localhost:8501/
- LangChain
- FAISS Vector Store
- Streamlit (for visualization)
- Ragas (for evaluation)
- Flexible integration with various LLMs (e.g., OpenAI GPT models, Groq, or others of your choice)
- Encoding both book content in chunks, chapter summaries generated by LLM, and quotes from the book.
- Anonymizing the question to create a general plan without biases or pre-trained knowledge of any LLM involved.
- Breaking down each task from the plan to be executed by custom functions with full control.
- Distilling retrieved content for better and accurate LLM generations, minimizing hallucinations.
- Answering a question based on context using a Chain of Thought, which includes both positive and negative examples, to arrive at a well-reasoned answer rather than just a straightforward response.
- Content verification and hallucination-free verification as suggested in "Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection" - https://arxiv.org/abs/2310.11511.
- Utilizing an ongoing updated plan made by an LLM to solve complicated questions. Some ideas are derived from "Plan-and-Solve Prompting" - https://arxiv.org/abs/2305.04091 and the "babyagi" project - https://github.com/yoheinakajima/babyagi.
- Evaluating the model's performance using
Ragasmetrics like answer correctness, faithfulness, relevancy, recall, and similarity to ensure high-quality answers.
Contributions are welcome! Please feel free to submit a pull request or open an issue for any suggestions or improvements.
Special thanks to Elad Levi for the valuable advice and ideas.
This project is licensed under the Apache-2.0 License - see the LICENSE file for details.
⭐️ If you find this repository helpful, please consider giving it a star!
Keywords: RAG, Retrieval-Augmented Generation, Agent, Langgraph, NLP, AI, Machine Learning, Information Retrieval, Natural Language Processing, LLM, Embeddings, Semantic Search
The Controllable RAG Agent is an advanced Retrieval-Augmented Generation (RAG) solution designed to tackle complex questions that simple semantic similarity-based retrieval cannot solve. It showcases a sophisticated deterministic graph acting as the "brain" of a highly controllable autonomous agent capable of answering non-trivial questions from your own data.
- Controllable RAG Agent: Deterministic graph-based reasoning, multi-step task decomposition, hallucination prevention, adaptive planning, verification and re-planning loops
- Traditional RAG: Single-step semantic similarity retrieval, direct answer generation without complex reasoning, no verification mechanisms
- LangChain RAG: Chain-based retrieval, simpler orchestration, less sophisticated reasoning
The Controllable RAG Agent handles questions requiring multi-hop reasoning and complex task decomposition.
- Sophisticated Deterministic Graph: Acts as the "brain" of the agent, enabling complex reasoning
- Controllable Autonomous Agent: Capable of answering non-trivial questions from custom datasets
- Hallucination Prevention: Ensures answers are solely based on provided data, avoiding AI hallucinations
- Multi-step Reasoning: Breaks down complex queries into manageable sub-tasks
- Adaptive Planning: Continuously updates its plan based on new information
- Performance Evaluation: Uses Ragas metrics for comprehensive quality assessment
- PDF Loading and Processing: Load PDF documents and split into chapters
- Text Preprocessing: Clean and preprocess text for better summarization and encoding
- Summarization: Generate extensive summaries using large language models
- Book Quotes Database Creation: Create database for specific questions needing quotes
- Vector Store Encoding: Encode content and summaries into vector stores for efficient retrieval
- Question Processing:
- Anonymize question by replacing named entities with variables
- Generate high-level plan for the anonymized question
- De-anonymize and break down into retrievable/answerable tasks
- Task Execution:
- Decide whether to retrieve information or answer based on context
- If retrieving: fetch from vector stores and distill
- If answering: generate response using chain-of-thought reasoning
- Verification and Re-planning:
- Verify content is grounded in original context
- Re-plan remaining steps based on new information
- Final Answer Generation: Produce answer using accumulated context and chain-of-thought
The solution is evaluated using Ragas metrics:
- Answer Correctness: Measures accuracy of the generated answer
- Faithfulness: Ensures answer is grounded in retrieved context
- Answer Relevancy: Measures how relevant the answer is to the question
- Context Recall: Measures coverage of relevant information
- Answer Similarity: Compares generated answer to reference
For questions like "How did the protagonist defeat the villain's assistant?", the agent:
- Identifies the protagonist
- Identifies the villain
- Identifies the villain's assistant
- Searches for confrontations between protagonist and villain
- Deduces the reason for defeating the assistant
This demonstrates sophisticated reasoning beyond simple retrieval.
The agent is built on LangGraph and supports:
- OpenAI: used by default (
gpt-4o) - Groq: wired up in
functions_for_pipeline.pyfor fast inference - Any LangChain-compatible provider: swap the chat model in
functions_for_pipeline.py
- Clone the repository
- Install dependencies:
pip install -r requirements.txt - Set up your LLM API keys (OpenAI, Anthropic, etc.)
- Prepare your PDF documents
- Run the agent with your custom questions
See the comprehensive guide on RAG techniques for additional context.
Yes! The agent is designed to work with any PDF documents. Load your custom documents, and the agent will:
- Process and split them appropriately
- Create vector stores for efficient retrieval
- Answer complex questions based solely on your data
The agent ensures answers are solely based on provided data through:
- Verification loops: Check generated content against original context
- Re-planning: Adjust approach if verification fails
- Grounded responses: Only use retrieved information, not pre-trained knowledge
Questions are anonymized by replacing named entities with variables. This helps:
- Generate more general plans
- Avoid bias from pre-trained knowledge about entities
- De-anonymize after planning to apply to specific context
- Agents Towards Production: Horizontal, code-first tutorials covering every step in building production-grade GenAI agents
- RAG Techniques: Comprehensive guide on RAG techniques
- GenAI Agents: Many other AI Agent implementations and tutorials
Apache 2.0 License - open-source, free to use commercially.
- GitHub Issues: github.com/NirDiamant/Controllable-RAG-Agent/issues
- Discord: Join our community
- Twitter: @NirDiamantAI
- LinkedIn: Connect
- Newsletter: DiamantAI Substack - 20,000+ subscribers



