01Responsibilities
Contribute to the IRIS Platform SRE and operations road map and execute planned research and development
Lead the design, deployment, and operations of large-scale systems on AWS (EKS) or Azure (AKS), ensuring reliability, scalability, and security.
Serve as the principal architect for platform reliability, performance, and disaster recovery strategies.
Troubleshoot active Production issues, provide RCA, co-ordinate with DevOps and Backend teams for permanent fixes. Primarily point of L2 contact for performance issues and flexible for non-office hours troubleshooting needs.
Able to install observability stack in Kubernetes such as EFK, Prometheus and Grafana
Able to create SLO/Grafana widgets using Prometheus queries. Should know all metric types and most common Prometheus functions.
Architect, install and govern Service Mesh (Istio/Linkerd) adoption for secure service-to-service communication, traffic management, and zero-trust networking
Apply SRE best practices, including SLIs, SLOs, error budgets, incident management, and post-mortem analysis.
Implement best alerting strategies to reduce noise and improve actionable incident detection
Optimize platform performance, scalability, and cloud cost efficiency (FinOps practices) across Kubernetes and cloud environments
Implement Infrastructure as Code using Terraform, CloudFormation, or AWS CDK for cloud and Kubernetes environments.
Optimize platform performance, scalability, and cost in cloud and hybrid environments.
Ensure compliance with security, governance, and operational standards across all deployments.
Mentor junior engineers and act as a technical authority on reliability, cloud architecture, and DevOps
Required Skills & Qualifications:
7+ years of experience in Site Reliability Engineering (SRE).
3+ years of hands-on experience working with Linux systems.
4+ years of commercial experience with Kubernetes.
2+ years of experience working with Docker.
4+ years of experience setting up and managing CI/CD pipelines.
4+ years of experience working with automation tools such as Terraform and Ansible.
Experience with containerization technologies, including Helm, and CI/CD pipelines.
Good knowledge of security best practices and vulnerability management tools (e.g., Acunetix, Snyk, CheckMarx, Trivy).
Experience troubleshooting production issues and performing root cause analysis.
Ability to work effectively in an Agile environment.
Preferred Skills & Qualifications:
Operational knowledge of databases (Postgres, ElasticSearch, Redis, or similar).
Exposure to configuring web servers such as Nginx.
Working knowledge of monitoring tools such as Grafana and Prometheus.
Working knowledge of a messaging framework such as Event Hub, Kafka, RabbitMQ, or similar.
Diversity & Inclusion Statement:
We are committed to building a diverse and inclusive team and encourage candidates from all backgrounds to apply.
About Us
SymphonyAI is building the leading enterprise AI SaaS company for digital transformation across the most critical and resilient growth industries, including retail, consumer packaged goods, financial crime prevention, manufacturing, media, and IT service management. Since its founding in 2017, SymphonyAI today serves 1500+