Skip to content
Srikanta Sahu

Projects07 Reliability

SRE Reliability Platform

SLIs, SLOs, and error budgets connected to burn-rate alerts and incident follow-up.

Problem

Telemetry existed, but it did not turn into reliability decisions. SLIs, error budgets, and incidents lived in different places.

Solution

Built a reliability platform that calculates SLIs, evaluates SLOs, tracks error budgets, detects burn-rate conditions, and ties an incident back to a reliability action.

Architecture

  1. Telemetry
  2. SLI
  3. SLO
  4. Error budget
  5. Burn rate
  6. Incident
  7. Reliability action

Technology

  • Prometheus
  • Grafana
  • Python
  • Kubernetes
  • OpenTelemetry

Engineering decisions

  • SLIs that measure the user-visible failure, not an easy internal counter.
  • Burn-rate alerts that fire early enough to act and late enough to stay quiet.
  • A clear owner for each service so an alert is not ownerless.
  • Joining an incident to the SLO it burned, not only to a ticket number.
  • A reliability report that a team can read without another slide deck.

Automation

Burn-rate conditions open the incident path and attach the SLO context. Post-incident actions are tracked against the same service.

Reliability

If the SLI query fails, the platform says the budget is unknown rather than drawing a green chart from empty data.

Security

Service ownership and alert routes are data, not a spreadsheet on someone's laptop. Access to budget views follows the same roles as the dashboards.