Projects06 Platform engineering
Observability-as-Code Platform
Application observability defined as configuration, then provisioned through Terraform, Helm, and GitOps.
Problem
Each application was onboarded to metrics, logs, traces, dashboards, and SLOs by hand, so the standard drifted and the next service started from scratch.
Solution
Defined application observability as configuration and automatically provisioned the collectors, dashboards, alerts, and SLO resources from that definition.
Architecture
- App definition
- CI validation
- Terraform / Helm
- Telemetry
- Dashboards and SLOs
- Deploy
Technology
- Terraform
- Kubernetes
- Helm
- Prometheus
- Grafana
- Loki
- OpenTelemetry
- GitHub Actions
Engineering decisions
- Reusable modules that still fit services with different signals.
- Detecting drift when someone changes a dashboard outside Git.
- Policy checks so an application cannot skip required alerts.
- A default onboarding path that is the same for every team.
- A GitOps workflow where Git remains the source of the live config.
Automation
A merged definition creates or updates collectors, dashboards, and alerts. Teams do not click those resources together in the UI.
Reliability
CI rejects a definition that is missing required fields. Apply is incremental, so a bad module does not rebuild the whole platform.
Security
Modules create scoped service accounts and keep datasource credentials in the secret store, not in the application repo.