01Overview
Description
We are seeking a Site Reliability Engineer (SRE) with 7+ years of experience to support and enhance the reliability, availability, and performance of critical banking systems at Truist. The role requires strong hands on expertise in cloud native platforms, observability, automation, and incident management, with a focus on reliability engineering and operational excellence
Required Technical Skills
Cloud & Infrastructure
Microsoft Azure
Kubernetes
OpenShift
Observability & Monitoring
Datadog
Dynatrace / AppDynamics
Splunk
Jenkins
Ansible
Automation & CI/CD
Python
Kafka
RabbitMQ
Exposure to Java and Node.js
Production Support & Incident Management
Strong experience handling major incidents
Production support in high availability, mission critical environments
Root cause analysis and reliability improvement
Engineer and enhance observability across systems and platforms
Define, implement, and track SLIs and SLOs
Design and build automation for recovery and self healing
Apply cloud native resiliency and failure isolation patterns
Lead major incident response with an engineering driven approach
Drive system level root cause fixes
Reduce long term incident volume through reliability engineering initiatives
Desired Skills
Analyze, optimize, and enable CI/CD pipelines to improve reliability outcomes
Supplementary Skills (Good to Have)
Advanced use of AIOps for predictive reliability insights
",
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying. .