Skip to content
Srikanta Sahu

Experience

Public-facing roles only. Internal systems, customer names, and unstated metrics are omitted.

  1. Current

    Morgan Stanley

    Senior Site Reliability Engineer

    Bengaluru · September 2021 — Present

    Lead enterprise observability and reliability initiatives for critical applications, partnering with engineering, monitoring, support, security, and vendor teams to improve system availability, scalability, and operational transparency.

    Impact

    • Built and operated an incident-response framework that contributed to a 25% year-over-year reduction in downtime.
    • Partnered with engineering teams to define performance-testing protocols for upgrades and patches, supporting 99.98% uptime for a critical application.
    • Automated routine maintenance work, reducing manual effort by 200 hours per month and reducing human error.
    • Supported high-volume workflows generating terabytes of data while maintaining uninterrupted processing.
    • Led L2 escalation activities, strengthened incident routing, and improved operational documentation and training.
    • Coordinated disaster-recovery and continuity testing, release-cycle component upgrades, and vulnerability remediation.
  2. Lowe's India

    Senior Software Engineer

    Bengaluru · September 2018 — September 2021

    Built and managed enterprise observability and monitoring platforms across on-premise, Kubernetes, and GCP environments.

    Impact

    • Architected monitoring using Prometheus, InfluxDB, and Grafana across on-premise, GCP, and Kubernetes platforms.
    • Built a unified observability solution for metrics, logs, and traces to improve operational visibility and DevOps efficiency.
    • Designed a highly available Prometheus and Grafana architecture with attention to storage, security, and reliability.
    • Onboarded and migrated 300+ microservices to metrics and logging, including golden-signal dashboards for teams.
    • Managed logging at terabyte-per-day scale and metrics and tracing across 250+ Kubernetes services.
    • Implemented distributed tracing to improve response-time analysis and application-performance visibility.
    • Led a proprietary-tool migration to open-source observability solutions, saving $1.2 million annually in licensing costs.

    Technologies

    • Prometheus
    • Grafana
    • InfluxDB
    • Kubernetes
    • GCP
  3. NTT DATA

    System Integration Engineer

    August 2016 — August 2018

    Supported monitoring automation and operational reliability for multiple client environments across UNIX, Windows, databases, web servers, and application servers.

    Impact

    • Configured and upgraded agent-based and agentless monitoring, and investigated daily monitoring and platform issues.
    • Gathered requirements and supported walkthroughs, testing, implementation, and documentation with client, vendor, and internal teams.
    • Managed monitoring incidents, requests, and changes using ITIL practices.
    • Provided monitoring-process guidance to L2 and L3 analysts and supported 24/7 server monitoring.

    Technologies

    • UNIX
    • Windows