Data Lakehouse Architecture
Free template — view it below, open it in draw.io, or customize it with AI in seconds.
The prompt behind this diagram
A modern data lakehouse architecture: sources (operational DBs, SaaS APIs, event streams), Airbyte ingestion, S3 raw/bronze/silver/gold zones with Delta Lake, Spark transformation jobs orchestrated by Airflow, dbt models, Snowflake serving layer, BI tools (Metabase) and a feature store.
Paste your own description (or Terraform / docker-compose / SQL schema) into draft1 and get a diagram like this for your exact system.
What this diagram shows
A data lakehouse architecture organises raw data into three medallion zones (Bronze, Silver, Gold) with transformation pipelines between them. Raw data flows from multiple sources into Bronze (immutable storage of all ingested data), then moves through Silver (cleansed, deduplicated, validated datasets) via Spark and dbt transformations, before reaching Gold (business-ready aggregates and dimensional tables). BI tools and analytics applications query the Gold layer. Delta Lake provides ACID transactions and schema governance across all zones, enabling data quality checks and lineage tracking throughout the pipeline.
Key components
- Data Sources — External systems, APIs, databases, and event streams that continuously feed raw data into the lakehouse.
- Bronze Zone — Immutable raw data landing area that preserves original data in its native format without transformation or filtering.
- Silver Zone — Intermediate layer where raw data is cleansed, deduplicated, validated, and structured using Spark and dbt transformations.
- Gold Zone — Refined, business-ready layer containing aggregated metrics, dimensional tables, and domain-specific datasets optimised for analytics.
- Delta Lake — Storage format and transaction layer providing ACID guarantees, schema enforcement, and data versioning across all zones.
- Spark and dbt — Transformation engines that orchestrate data movement between zones, apply business logic, and maintain data quality.
- BI and Analytics Layer — Query tools, dashboards, and applications that consume Gold zone data for reporting and analysis.
When to use it
Use this template when building modern data platforms that must handle high-volume ingestion, require data quality enforcement, and need both historical archives and real-time analytics. It suits organisations transitioning from data warehouses to lakehouses, those running on Databricks, Delta Lake, or Apache Spark, and teams needing clear separation between raw data preservation and curated analytics. Best for companies with mature data engineering and analytics practices.
Common mistakes
- Treating Bronze as a dumping ground without tracking schema or data quality, leading to unusable raw data and broken downstream pipelines.
- Skipping the Silver layer and transforming raw data directly to Gold, which removes audit trails and makes debugging data issues nearly impossible.
- Overloading the Gold layer with every possible metric and dimension instead of focusing on specific business domains, causing query performance to degrade and maintenance burden to increase.
Adapting it to your system
Replace Bronze/Silver/Gold labels with your specific naming convention if your organisation uses different terms. Swap Spark and dbt for your actual transformation tools (Airflow, Prefect, Flink, or native cloud services). Add your real data sources on the left (Salesforce, Kafka, databases, cloud storage). Insert your actual BI tools on the right (Tableau, Looker, PowerBI). Replace Delta Lake with your chosen lakehouse format if using Iceberg or Hudi. Adjust the number of transformation steps based on your data quality requirements and business complexity.
More templates
AWS VPC Multi-AZ Architecture
A production AWS VPC layout template: public/private/data subnets across two AZs with NAT, RDS multi-AZ and S3 endpoin
AWS EKS Cluster Architecture
An EKS reference template: control plane, node groups, ALB ingress, ECR, IAM roles for service accounts and storage.
AWS ECS Fargate Architecture
Serverless containers on AWS: ALB, Fargate services, SQS decoupling, RDS and Redis — a production ECS template.
Azure 3-Tier Web Architecture
The Azure counterpart of the classic 3-tier stack: Front Door, App Gateway, App Services, SQL and Redis in a VNet.
GCP Web Application Architecture
A serverless GCP stack template: Cloud Run, Cloud SQL, Memorystore, Pub/Sub and CDN-fronted load balancing.
Kafka Event Streaming Pipeline
End-to-end event streaming: CDC ingestion, a three-broker cluster, stream processing and analytical sinks.
ML Training & Inference Pipeline
MLOps reference template: feature store, tracked training, registry, real-time + batch inference and drift-driven retr