Current
Morgan Stanley
Senior Site Reliability Engineer
Bengaluru · September 2021 — Present
Lead enterprise observability and reliability initiatives for critical applications, partnering with engineering, monitoring, support, security, and vendor teams to improve system availability, scalability, and operational transparency.
Impact
- Built and operated an incident-response framework that contributed to a 25% year-over-year reduction in downtime.
- Partnered with engineering teams to define performance-testing protocols for upgrades and patches, supporting 99.98% uptime for a critical application.
- Automated routine maintenance work, reducing manual effort by 200 hours per month and reducing human error.
- Supported high-volume workflows generating terabytes of data while maintaining uninterrupted processing.
- Led L2 escalation activities, strengthened incident routing, and improved operational documentation and training.
- Coordinated disaster-recovery and continuity testing, release-cycle component upgrades, and vulnerability remediation.