Location: Guwahati
Experience: 8-10 Years
Job Band: E4
Role Type: Individual Contributor (IC)
Travel: No
Reports To: Lead Site Reliability Engineer
We are looking for an experienced Senior Site Reliability Engineer to help build and maintain highly reliable, scalable, secure, and resilient platforms and infrastructure. The role involves troubleshooting complex production issues, improving system performance, implementing automation, and ensuring the reliability and availability of critical services.
Build and promote a strong Site Reliability Engineering (SRE) culture by sharing best practices, documentation, approaches, and code across engineering teams.
Automate manual and repetitive tasks using software, scripting, and engineering best practices.
Troubleshoot complex, cross-platform issues involving Operating Systems, Networking, Databases, and Cloud-based SaaS environments.
Handle and resolve live production incidents and ensure timely restoration of services.
Monitor application and infrastructure performance and implement improvements to enhance stability, availability, and performance.
Conduct system analysis and configuration management to improve system reliability and efficiency.
Design, develop, and deploy software and systems that improve observability, product reliability, and operational efficiency.
Manage and monitor the deployment and orchestration of servers, Docker containers, databases, and backend infrastructure.
Develop and maintain Runbooks and Standard Operating Procedures (SOPs) for recurring production issues.
Perform regular Incident Analysis / Root Cause Analysis (RCA) and implement long-term solutions to prevent recurring incidents.
Collaborate closely with Product, Development, Business Analysis, and Incident Management teams to ensure reliable service delivery.
Strong experience in Site Reliability Engineering, Production Support, and Infrastructure Monitoring.
Hands-on experience with infrastructure performance monitoring and analysis using standard monitoring tools.
Strong hands-on experience with Docker and Kubernetes.
Experience with Infrastructure as Code (IaC) tools such as:
Terraform
CloudFormation
Ansible
Hands-on experience with large-scale databases and distributed technologies.
Good knowledge and hands-on experience with Apache Kafka / Confluent Kafka Platform.
Basic to intermediate programming and scripting skills.
Strong troubleshooting and problem-solving skills across OS, networking, databases, cloud, and application infrastructure.
Experience handling production incidents, RCA, observability, and system reliability improvements.
B.Tech / B.E. in Computer Science, Information Technology, or a related technical discipline is preferred.
8-10 years of overall experience, preferably with significant experience in:
Site Reliability Engineering
DevOps / Cloud Infrastructure
Production Support
Infrastructure & Application Monitoring
Incident Management
Automation and Observability
Internal Stakeholders:
Product Manager Digital Factory
Business Analyst BTG Team
Incident Management Team
Development Team
The ideal candidate should have strong experience in SRE/DevOps, cloud infrastructure, containerization, Kubernetes, IaC, monitoring, distributed systems, and production incident management, along with the ability to identify root causes and drive permanent solutions.