Observability Stack Architecture
Free template — view it below, open it in draw.io, or customize it with AI in seconds.
The prompt behind this diagram
An observability architecture: applications emitting OpenTelemetry traces, metrics and logs; OTel collector; Prometheus for metrics with Alertmanager to PagerDuty; Loki for logs; Tempo for traces; Grafana dashboards on top; SLO burn-rate alerts.
Paste your own description (or Terraform / docker-compose / SQL schema) into draft1 and get a diagram like this for your exact system.
What this diagram shows
This diagram represents the flow of observability signals (metrics, logs, traces) from instrumented applications through collection and storage layers into a unified dashboard and alerting system. Applications emit telemetry using OpenTelemetry SDKs or agents, which route metrics to Prometheus, logs to Loki, and traces to Tempo. These three backends store their respective signals independently. Prometheus evaluates alert rules and fires notifications, while Grafana queries all three data sources, dashboards surface service-level objectives (SLOs) calculated from the stored data, and operators respond to firing alerts. The architecture separates concerns by data type while maintaining single-pane-of-glass visibility.
Key components
- Application Instrumentation — Emits metrics, logs, and spans using OpenTelemetry SDKs, libraries, or auto-instrumentation agents embedded in services.
- OpenTelemetry Collector — Receives signals from instrumented applications via OTLP protocol, applies transformations and filtering, then exports to backend storage systems.
- Prometheus — Scrapes or receives metrics via push, stores time-series data, and evaluates alert rules against metric thresholds to generate alerts.
- Loki — Stores application and infrastructure logs indexed by labels, allowing log queries and filtering without full-text index overhead.
- Tempo — Stores distributed traces with span relationships and timing information, enabling service dependency mapping and latency analysis.
- Grafana — Queries Prometheus, Loki, and Tempo simultaneously to surface metrics, logs, traces, and SLO compliance in interactive dashboards.
- Alerting and Notification — Receives alerts fired by Prometheus or Grafana alert rules and routes them to incident management systems, chat platforms, or PagerDuty.
When to use it
Use this pattern when you need to correlate metrics, logs, and traces across distributed systems without vendor lock-in. It works well for teams already running Kubernetes or container infrastructure, adopting OpenTelemetry standards, and wanting to avoid separate point solutions for each signal type. Choose it when SLO monitoring and alert-driven incident response are priorities, and your team has capacity to operate multiple stateful backends.
Common mistakes
- Treating OpenTelemetry Collector as optional and having each application push directly to backends, which bypasses crucial signal processing, sampling, and reliability.
- Storing all logs indefinitely in Loki without retention policies or partitioning strategies, causing unbounded disk growth and expensive query latency.
- Configuring alert rules only on metrics without correlating them to trace data or logs, leading to alerts firing without enough context to troubleshoot root cause.
Adapting it to your system
Replace the generic application box with your actual services: microservices, Lambda functions, or containerized workloads. Adjust OpenTelemetry Collector configuration for your environment (Kubernetes DaemonSet, sidecar, or gateway mode). Specify Prometheus scrape targets matching your infrastructure. Define Loki label schema around your deployment boundaries (namespace, service, environment). Add storage backends appropriate to your scale: S3 for Loki, object storage for Tempo, or local PVs for Prometheus in non-production. Tailor SLO queries to your actual error budgets and business objectives.
More templates
AWS VPC Multi-AZ Architecture
A production AWS VPC layout template: public/private/data subnets across two AZs with NAT, RDS multi-AZ and S3 endpoin
AWS EKS Cluster Architecture
An EKS reference template: control plane, node groups, ALB ingress, ECR, IAM roles for service accounts and storage.
AWS ECS Fargate Architecture
Serverless containers on AWS: ALB, Fargate services, SQS decoupling, RDS and Redis — a production ECS template.
Azure 3-Tier Web Architecture
The Azure counterpart of the classic 3-tier stack: Front Door, App Gateway, App Services, SQL and Redis in a VNet.
GCP Web Application Architecture
A serverless GCP stack template: Cloud Run, Cloud SQL, Memorystore, Pub/Sub and CDN-fronted load balancing.
Kafka Event Streaming Pipeline
End-to-end event streaming: CDC ingestion, a three-broker cluster, stream processing and analytical sinks.
Data Lakehouse Architecture
Bronze/silver/gold lakehouse template: ingestion, Delta Lake zones, Spark + dbt transforms and a BI serving layer.