ML Training & Inference Pipeline
Free template — view it below, open it in draw.io, or customize it with AI in seconds.
The prompt behind this diagram
A machine learning pipeline architecture: feature store, training pipeline with experiment tracking (MLflow), model registry, batch inference job, real-time inference API behind a load balancer with autoscaling, monitoring with drift detection, retraining trigger loop.
Paste your own description (or Terraform / docker-compose / SQL schema) into draft1 and get a diagram like this for your exact system.
What this diagram shows
A complete machine learning lifecycle from raw data to production predictions. Data flows through a feature store that prepares and versions inputs, into a training service that logs metrics and model artifacts, which are registered and versioned in a model registry. From there, the trained model routes to both batch inference (for bulk scoring) and real-time inference (for on-demand predictions). A monitoring component tracks prediction drift and data drift, triggering automated retraining when thresholds breach. Feedback loops from inference results feed back to monitoring and feature engineering stages.
Key components
- Feature Store — Centralises, versions, and serves engineered features consistently to training and inference workloads.
- Training Service — Executes model training runs, logs hyperparameters and metrics, and outputs candidate models.
- Model Registry — Stores, versions, stages (dev/staging/prod), and tracks metadata for all trained models.
- Batch Inference — Scores large datasets on a schedule, writing predictions to a data warehouse or lakehouse.
- Real-time Inference API — Serves low-latency predictions to applications, retrieving features on-demand from the feature store.
- Monitoring & Drift Detection — Measures prediction drift, data drift, and model performance; emits alerts or triggers retraining when anomalies occur.
- Retraining Orchestrator — Automatically re-runs training pipelines when drift thresholds breach or on a scheduled cadence.
When to use it
Use this template when building end-to-end ML systems that require reproducibility, model governance, and automated model updates in production. It suits organisations operating multiple models in parallel, needing audit trails for compliance, or supporting both batch and real-time prediction workloads. Typical use cases include fraud detection, recommendation engines, forecasting, and demand prediction where model decay is expected and retraining is frequent.
Common mistakes
- Omitting the feature store and treating features as ad-hoc transformations, leading to inconsistency between training and serving environments.
- Assuming a single inference path (batch only or real-time only) when production often demands both asynchronous scoring and synchronous API calls.
- Skipping drift monitoring and retraining triggers, so stale models continue serving predictions even after performance degrades in production.
Adapting it to your system
Identify your raw data sources (databases, streams, data lake) and map them into feature store tables. Name the specific ML framework (scikit-learn, XGBoost, PyTorch) and training orchestrator (Airflow, Kubeflow, Prefect). Choose your inference deployment pattern (FastAPI containers, AWS SageMaker endpoints, Seldon) and batch runner (Spark, dbt, Lambda functions). Define drift metrics relevant to your use case (prediction distribution shift, feature value changes, business KPIs). Connect your monitoring tool (Datadog, Prometheus, custom dashboards) and alert channels to your retraining trigger logic.
More templates
AWS VPC Multi-AZ Architecture
A production AWS VPC layout template: public/private/data subnets across two AZs with NAT, RDS multi-AZ and S3 endpoin
AWS EKS Cluster Architecture
An EKS reference template: control plane, node groups, ALB ingress, ECR, IAM roles for service accounts and storage.
AWS ECS Fargate Architecture
Serverless containers on AWS: ALB, Fargate services, SQS decoupling, RDS and Redis — a production ECS template.
Azure 3-Tier Web Architecture
The Azure counterpart of the classic 3-tier stack: Front Door, App Gateway, App Services, SQL and Redis in a VNet.
GCP Web Application Architecture
A serverless GCP stack template: Cloud Run, Cloud SQL, Memorystore, Pub/Sub and CDN-fronted load balancing.
Kafka Event Streaming Pipeline
End-to-end event streaming: CDC ingestion, a three-broker cluster, stream processing and analytical sinks.
Data Lakehouse Architecture
Bronze/silver/gold lakehouse template: ingestion, Delta Lake zones, Spark + dbt transforms and a BI serving layer.