01Key Responsibilities
Proficient in Splunk/ELK, and Datadog.
Experience with observability tools such as Prometheus/InfluxDB, and Grafana.
Possesses strong knowledge of at least one scripting language such as Python, Bash, Powershell or any other relevant languages.
Design, develop, and maintain observability tools and infrastructure.
Collaborate with other teams to ensure observability best practices are followed.
Develop and maintain dashboards and alerts for monitoring system health.
Troubleshoot and resolve issues related to observability tools and infrastructure.
Engage in and improve the lifecycle of services from conception to EOL, including: system design consulting, and capacity planning
Define and implement standards and best practices related to: System Architecture, Service delivery, metrics and the automation of operational tasks
Support services, product & engineering teams by providing common tooling and frameworks to deliver increased availability and improved incident response.
Improve system performance, application delivery and efficiency through automation, process refinement, postmortem reviews, and in-depth configuration analysis
Collaborate closely with engineering professionals within the organization to deliver reliable services
Identify and eliminate operational toil by treating operational challenges as a software engineering problem
Actively participate in incident response, including on-call responsibilities
Partner with stakeholders to influence and help drive the best possible technical and business outcomes
Guide junior team members and serve as a champion for Site Reliability Engineering
Engineering degree, or a related technical discipline, and 10+years of experience in SRE.
Experience coding in higher-level languages (e.g., Python, Javascript, C++, or Java)
Knowledge of Cloud based applications & Containerization Technologies
Demonstrated understanding of best practices in metric generation and collection, log aggregation pipelines, time-series databases, and distributed tracing
Ability to analyze current technology utilized and engineering practices within the company and develop steps and processes to improve and expand upon them
Working experience with industry standards like Terraform, Ansible.
(Experience, Education, Certification, License and Training)
Must have hands-on experience working within Engineering or Cloud.
Experience with public cloud platforms (e.g. GCP, AWS, Azure)
Experience in configuration and maintenance of applications & systems
infrastructure. Experience with distributed system design and architecture
Experience building and managing CI/CD Pipelines
Company Overview:
UKG is the Workforce Operating Platform that puts workforce understanding to work. With the world's largest collection of workforce insights, and people-first AI, our ability to reveal unseen ways to build trust, amplify productivity, and empower talent, is unmatched. It's this expertise that equips our customers with the intelligence to solve any challenge in any industry because great organizations know their workforce is their competitive edge. Learn more at ukg.com.
UKG is proud to be an equal opportunity employer and is committed to promoting diversity and inclusion in the workplace, including the recruitment process.
Disability Accommodation in the Application and Interview Process
For individuals with .