RAG Chatbot Architecture
Free template — view it below, open it in draw.io, or customize it with AI in seconds.
The prompt behind this diagram
A production RAG chatbot architecture: document sources, ingestion pipeline with chunking and embedding, vector database (Qdrant), query service with hybrid search and reranker, LLM API (Claude), response with citations, feedback loop, evaluation harness.
Paste your own description (or Terraform / docker-compose / SQL schema) into draft1 and get a diagram like this for your exact system.
What this diagram shows
A RAG chatbot architecture shows how a conversational system retrieves relevant documents, ranks them by relevance, and synthesizes responses using an LLM. The flow begins with user input, moves through vector search against an embedded document store, applies reranking to filter low-quality results, passes context to the LLM with a prompt, and outputs a response. Optional feedback loops evaluate answer quality and update embeddings. This design separates retrieval (finding potentially relevant documents) from ranking (keeping only the best ones) from generation (composing the actual response), allowing each stage to be tuned independently.
Key components
- User Query Input — Entry point where the user submits a question or prompt to the chatbot system.
- Vector Embedding Service — Converts the user query into a high-dimensional vector using a pre-trained embedding model so it can be compared against stored document embeddings.
- Vector Database — Stores pre-computed embeddings of ingested documents (chunks) and performs fast similarity search to retrieve candidate documents.
- Reranker Model — Takes the candidate documents from vector search and scores them using cross-attention, keeping only the top-k most relevant results to reduce noise.
- LLM with Prompt Template — Generates the final response by processing the reranked context chunks plus the original query through a structured prompt.
- Response Output — Delivers the LLM-generated answer to the user.
- Evaluation and Feedback — Logs query-response pairs and optionally gathers user feedback to measure answer quality and trigger retraining or re-indexing.
When to use it
Use this architecture when you need a chatbot that must ground answers in a specific knowledge base (documentation, research papers, product data) rather than relying solely on an LLM's training data. It is essential when accuracy, attribution, and freshness matter more than speed alone. Choose it for question-answering systems over internal wikis, customer support bots, or domain-specific assistants where hallucination or outdated information is costly.
Common mistakes
- Skipping the reranking step and passing all vector search results to the LLM, which wastes context window and degrades response quality by mixing weak matches with strong ones.
- Treating vector embeddings as a solved problem and never updating them after document ingestion, so new or revised documents are never discoverable in search.
- Failing to log failed queries and user feedback, making it impossible to identify which document chunks are missing or which retrieval failures cost users the most.
Adapting it to your system
Replace the embedding model with one tuned to your domain (e.g., BioGPT for medical text or CodeBERT for repositories). Adjust the vector database choice based on scale: Pinecone for managed simplicity, Qdrant or Milvus for self-hosted control. Swap the reranker for your latency budget; cross-encoder models are slower but more accurate than lightweight ranking heuristics. Modify the prompt template to include your domain's terminology and output format requirements. Connect the feedback loop to your preferred analytics tool to surface which queries are failing most often.
More templates
AWS VPC Multi-AZ Architecture
A production AWS VPC layout template: public/private/data subnets across two AZs with NAT, RDS multi-AZ and S3 endpoin
AWS EKS Cluster Architecture
An EKS reference template: control plane, node groups, ALB ingress, ECR, IAM roles for service accounts and storage.
AWS ECS Fargate Architecture
Serverless containers on AWS: ALB, Fargate services, SQS decoupling, RDS and Redis — a production ECS template.
Azure 3-Tier Web Architecture
The Azure counterpart of the classic 3-tier stack: Front Door, App Gateway, App Services, SQL and Redis in a VNet.
GCP Web Application Architecture
A serverless GCP stack template: Cloud Run, Cloud SQL, Memorystore, Pub/Sub and CDN-fronted load balancing.
Kafka Event Streaming Pipeline
End-to-end event streaming: CDC ingestion, a three-broker cluster, stream processing and analytical sinks.
Data Lakehouse Architecture
Bronze/silver/gold lakehouse template: ingestion, Delta Lake zones, Spark + dbt transforms and a BI serving layer.