Projects03 AI / SRE
AI SRE / Autonomous Incident Investigation
An investigation assistant that queries metrics, logs, cluster state, and runbooks, then drafts an explainable RCA for a human to approve.
Problem
When an incident fires, an engineer still has to query metrics, logs, cluster state, and runbooks by hand, then rebuild a timeline before they can recommend a fix. That is slow and easy to get wrong.
Solution
Built an investigation assistant that queries observability and infrastructure systems, correlates evidence, consults runbooks, constructs a timeline, and produces an explainable root-cause write-up with a remediation recommendation. A human still approves any change. The assistant does not remediate on its own.
Architecture
- Incident
- Investigation
- Prometheus
- Loki
- Kubernetes
- Runbook RAG
- Timeline
- Root cause
- Human approval
Technology
- Prometheus
- Loki
- OpenTelemetry
- Kubernetes
- Python
- LLM
- RAG
- MCP
- n8n
Engineering decisions
- Tool access with narrow permissions so the assistant can query without changing the system.
- Evidence-based RCA: every claim has to point at a metric, log, event, or runbook passage.
- Correlating metrics, logs, and cluster events onto one timeline.
- Keeping a human in the loop so a recommendation never becomes an unreviewed change.
Automation
The first investigation pass — gathering signals, assembling a timeline, and drafting the RCA — no longer starts as a blank page.
Reliability
If a tool call fails, the write-up says which evidence is missing instead of guessing. The path stops at a recommendation.
Security
Credentials stay on the server. The assistant can read the signals it is allowed to see and cannot apply a change without approval.
AI layer
RAG retrieves runbooks and postmortems. MCP exposes metrics, logs, and cluster state as tools. The model proposes; a person decides.