01Key Responsibilities
Support and maintain Linux-based servers and telephony services in production.
Investigate and resolve incidents in a high-load, distributed environment.
Participate in on-call shifts and ensure the stability of systems under strict SLAs.
Analyze service performance, reliability, and architecture bottlenecks; propose improvements.
Work with development teams to safely deliver and validate changes before production deployment.
Contribute ideas and help evolve team processes, automation, and monitoring practices.
02Requirements
8+ years of Strong hands-on experience with UNIX/Linux systems, with deep proficiency in CLI-based diagnostics, system administration, and production troubleshooting at scale.
Good understanding of networking protocols or SIP.
Strong hands-on experience with Kubernetes (k8s) and containerized environments.
Proven track record of working in production environments, with a careful and methodical approach to changes (testing before deployment, rollback planning, risk mitigation).
Understanding of high-availability systems, fault tolerance, and performance optimization.
Experience automating tasks with Python, Golang, or Shell scripts.
Mindset of an SRE: you treat operations as an engineering discipline and continuously look for ways to make systems more reliable and efficient.
Good command of English (B2 or higher) ability to communicate effectively with distributed international teams (both written and spoken).
Would be a plus:
Deep expertise in one or more areas (please highlight your strengths in your application).
Hands-on experience with Kamailio, Apache Kafka, Nginx, ZeroMQ.
Experience with AWS/EKS, Terraform, and Ansible for deployment and infrastructure automation.
Experience with CI/CD pipelines (e.g., GitLab CI, Jenkins, ArgoCD).
Knowledge of monitoring stacks like Zabbix, TICK, ELK, Grafana.
What youll get:
Work in a strong, experienced SRE team that maintains global infrastructure across multiple regions.
Hands-on experience debugging Java and C++ applications in large distributed systems (Kafka, Zookeeper, Kamailio, Nginx, etc.).
Opportunity to influence how the team works your ideas for tools, automation, or process improvements will be heard and implemented.
Real experience achieving five-nines availability (99.999%) in production.
Continuous learning, complex technical challenges, and a supportive environment. .