01Responsibilities
API Gateway Platform Management
Own and operate enterprise API Gateway platforms, including:
Apigee (SaaS and OPDK)
Azure API Management (APIM)
AWS API Gateway
Ensure high availability, fault tolerance, and scalability of API gateway infrastructure across regions and environments
Manage API lifecycle concerns including security, traffic management, throttling, versioning, and observability
Design, develop, and deploy AI-powered solutions using no-code, low-code, and advanced platforms, translating business needs into scalable applications that enhance products, workflows, and decision-making
Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Cloud Infrastructure & AMI Management
Manage API gateway infrastructure hosted on AWS and Azure, with solid ownership of:
AMI creation, rotation, and patching
Zero downtime upgrades and rolling deployments
Ensure infrastructure meets security, resiliency, and compliance standards
Use Infrastructure as Code (Terraform, Ansible) as the default for all platform changes
AI First Operations & Automation
Treat automation and AI as the first option, not an afterthought
Apply AI driven and AIOps capabilities to:
Detect anomalies and traffic issues across API gateways
Predict failures and capacity risks
Reduce alert noise and accelerate root cause analysis
Build self healing mechanisms and automated runbooks for common API gateway failure scenarios
Continuously eliminate manual operational work through intelligent automation
Production Support & Incident Ownership
Independently own ServiceNow P1 and P2 incidents related to API gateway platforms
Lead incident triage, communication, and resolution during high severity outages
Perform deep root cause analysis and ensure permanent fixes through automation and platform improvements
Participate in on call rotations and act as an escalation point for API platform issues
Observability & Reliability
Design and operate observability for API gateways using:
OpenSearch
Dynatrace
CloudWatch
Splunk
Track and improve SLIs, SLOs, latency, error rates, and traffic patterns for APIs
Use data and AI insights to proactively improve platform reliability
Security & Access Management
Ensure solid authentication and authorization for API gateways (IAM, OAuth, mTLS, secrets management) .