01Overview
Technology
- 1003
Job Title: Principal Site Reliability Engineer (SRE), Digital Business
Location: Mumbai
Employment Type: Full time
About Us
Sonyliv is a leading OTT platform revolutionizing the way audiences consume entertainment. With millions of users across the globe, our mission is to deliver seamless, high-quality, and reliable streaming experiences. We are looking for a Principal Site Reliability Engineer (SRE) to join our team and take ownership of ensuring the availability, scalability, and performance of our critical systems.
Job Summary
Key Responsibilities
- Development & SRE Mindset: Leverage your experience as a developer and SRE to build tools, automation, and systems to improve system reliability and operational efficiency.
- Incident Management: Respond to and resolve critical system issues promptly, including being available for on-call support and handling emergencies during non-business hours, including late nights.
- Infrastructure Management: Design, deploy, and manage infrastructure solutions using containers (Docker/Kubernetes), networks, and CDNs to ensure scalability and performance.
- Observability: Drive best practices in observability, including metrics, logging, and tracing, to enhance system monitoring and proactive issue resolution. Implement and maintain observability tools like Prometheus, Grafana, ELK stack, or DataDog.
- Reliability and Performance: Proactively identify areas for improvement in system reliability, performance, and scalability, and define strategies and best practices to address them.
- Collaboration and Communication: Work closely with cross-functional teams, including development, QA, and support, to align goals and improve operational excellence. Communicate effectively across teams and stakeholders.
- CI/CD and Automation: Build and enhance CI/CD pipelines to improve deployment reliability and efficiency. Automate repetitive tasks and processes wherever possible.
Required Skills and Experience
- Experi .