01Key Responsibilities
Service Delivery & Operations
- Own end-to-end service delivery for multiple critical applications, ensuring high availability, stability, and performance across production environments.
- Lead 24/7 application support operations (L2/L3), including on-call rotations, incident bridges, escalations, and stakeholder communications.
- Act as the primary escalation point for major incidents, driving resolution until closure and RCA sign-off.
- Ensure consistent adherence to SLAs, OLAs, and KPIs, with continuous tracking and reporting.
SRE & Reliability Engineering
- Apply SRE principles to application support, including:
- Definition and tracking of SLIs, SLOs, and Error Budgets
- Improving reliability, scalability, and fault tolerance
- Drive initiatives to reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through automation, improved alerting, and runbooks.
- Partner with engineering teams to eliminate toil and improve operational efficiency.
Monitoring, Observability & AIOps
- Lead the design and adoption of monitoring and observability platforms using tools such as:
- Prometheus, Grafana, ELK, Dynatrace, OpenTelemetry (OTel)
- Ensure end-to-end visibility across infrastructure, applications, and business transactions.
- Implement proactive monitoring, intelligent alerting, and anomaly detection to prevent incidents before business impact.
- Drive adoption of AIOps and automated incident management (auto-remediation, automated runbooks where applicable).
Incident, Problem & Change Management
- Lead Major Incident Management (MIM) including war rooms, stakeholder updates, and executive reporting.
- Ensure timely and high-quality Root Cause Analysis (RCA) with preventive action plans.
- Govern problem management to identify recurring issues and drive long-term fixes.
- Oversee change and release management, ensuring minimal risk to production systems.
Stakeholder & Vendor Management
- Act as a trusted partner for business stakeholders, engineering teams, vendors, and clients.
- Communicate service health, risks, and improvement plans to senior leadership.
- Manage third-party vendors and support partners, ensuring contract compliance and service quality.
Continuous Improvement & Transformation
- Drive operational excellence initiatives to improve uptime, performance, and customer satisfaction.
- Lead transformation programs involving cloud migration, DevOps, and SRE adoption.
- Identify automation opportunities to reduce manual effort and operational cost.
- Support DR planning, testing, and compliance with RTO/RPO requirements.
Required Experience & Skills
Experience
- 15+ years of experience in Application Support / Production Operations, with 5+ years in a Service Delivery Manager / SRE / Operations Leadership role.
- Proven experience managing large application portfolios (50+ applications) in enterprise environments.
- Strong background in banking, financial services, or large regulated enterprises .