01Key Responsibilities
Hands-On Engineering & Troubleshooting
Programming and debugging .NET/Java applications to support reliability and performance goals, showcasing strong problem-solving abilities.
Diagnose and resolve complex application issues within production environments, including root cause analysis (RCA) and performance debugging.
Monitoring, Observability & Analysis
Maintain, aggregate and analyze logs from observability tools such as ELK Stack, Application Insights, Splunk & Kibana to monitor application health, performance bottlenecks and user experience.
Automation
Maintain automation scripts (PowerShell, Python) to streamline monitoring and operational tasks.
Identify repetitive manual tasks and automate them to improve operational efficiency.
Advocate for and implement process improvements and tooling enhancements.
Incident Response & Reliability
Deep understanding of SRE principles, including SLAs, SLOs, error budgets and reliability-centric system design.
Lead incident management, postmortems and implement preventative measures.
Infrastructure
Perform IIS administration and environment setup on Windows.
Manage applications in both on-premises and cloud (Azure) environments.
Set up and maintain environments with a focus on application support while being exposed to infrastructure knowledge as needed.
Collaboration & Mentorship
Navigate easily across Chubbs other departments to foster collaboration and issue resolution.
Share knowledge and mentor junior SREs, promoting best practices and a culture of reliability.
Qualifications:
10-15 years of hands-on SRE or application support engineering experience, with strong production support exposure.
Proficiency in PowerShell scripting, Python programming and .NET/Java development/debugging with a focus on automation, tooling and system integration.
Strong reasoning, analytical thinking and troubleshooting skills for applications, including RCA & memory debugging.
Experience with observability and monitoring tools (e.g., ELK Stack, Application Insights, Splunk, Kibana, AppDynamics, DynaTrace).
Basic/intermediate knowledge of databases such as MS SQL Server.
Strong communication skills and ability to work under pressure.
Nice to Have:
Experience with AI technologies (e.g. Claude) and their application in enhancing system reliability and performance.
Experience with infrastructure SRE roles.
Background in regulated or high-compliance industries.
Familiarity with chaos engineering, performance optimization or fault injection.
Familiarity with Azure cloud infrastructure and services (e.g. PaaS and identity management such as Active Directory, Azure AD).
Soft Skills:
Proactive, detail-oriented, and able to handle production-critical issues.
Collaborative mindset and willingness to mentor others.
What Success Looks Like
Rapid, effective resolution of incidents and performance .