01Overview
Years of Experience 58 YearsDesign, implement, and manage highly available, scalable, and secure cloud infrastructure on AWS.Build and maintain an end-to-end observability platform using OpenTelemetry, Grafana, Datadog, CloudWatch, and related tools.Implement AIOps capabilities, including:LLM-assisted incident triageAI-powered root cause analysisML-driven forecasting and anomaly detectionIntelligent alert correlation and noise reductionLead production incident management, on-call response, postmortems, and Root Cause Analysis (RCA).Automate operational workflows using Infrastructure as Code (Terraform/CloudFormation) and CI/CD pipelines.Drive infrastructure rightsizing, capacity planning, utilization analysis, and cloud cost optimization.Build dashboards, SLOs, SLIs, and error budgets to improve service reliability.Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks.Monitor application and infrastructure health while ensuring high uptime and service performance. .