01Key Responsibilities
Hands-On Engineering & Troubleshooting
- Programming and debugging .NET/Java applications to support reliability and performance goals, showcasing strong problem-solving abilities.
- Diagnose and resolve complex application issues within production environments, including root cause analysis (RCA) and performance debugging.
Monitoring, Observability & Analysis
- Maintain, aggregate and analyze logs from observability tools such as ELK Stack, Application Insights, Splunk & Kibana to monitor application health, performance bottlenecks and user experience.
Automation
- Maintain automation scripts (PowerShell, Python) to streamline monitoring and operational tasks.
- Identify repetitive manual tasks and automate them to improve operational efficiency.
- Advocate for and implement process improvements and tooling enhancements.
Incident Response & Reliability
- Deep understanding of SRE principles, including SLAs, SLOs, error budgets and reliability-centric system design.
- Lead incident management, postmortems and implement preventative measures.
Infrastructure
- Perform IIS administration and environment setup on Windows.
- Manage applications in both on-premises and cloud (Azure) environments.
- Set up and maintain environments with a focus on application support while being exposed to infrastructure knowledge as needed.
Collaboration & Mentorship
- Navigate easily across Chubbs other departments to foster collaboration and issue resolution.
- Share knowledge and mentor junior SREs, promoting best practices and a culture of reliability.
Qualifications:
- 5-10 years of hands-on SRE or application support engineering experience, with strong production support exposure.
- Proficiency in PowerShell scripting, Python programming and .NET/Java development/debugging with a focus on automation, tooling and system integration.
- Strong reasoning, analytical thinking and troubleshooting skills for applications, including RCA & memory debugging.
- Experience with observability and monitoring tools (e.g., ELK Stack, Application Insights, Splunk, Kibana, AppDynamics, DynaTrace).
- Basic/intermediate knowledge of databases such as MS SQL Server.
- Strong communication skills and ability to work under pressure.
Nice to Have:
- Experience with AI technologies (e.g. Claude) and their application in enhancing system reliability and performance.
- Experience with infrastructure SRE roles.
- Background in regulated or high-compliance industries.
- Familiarity with chaos engineering, performance optimization or fault injection.
- Familiarity with Azure cloud infrastructure and services (e.g. PaaS and identity management such as Active Directory, Azure AD).
Soft Skills:
- Proactive, detail-oriented, and able to handle production-critical issues.
- Collaborative mindset and willingness to mentor others.
What Success Looks Like
- Rapid, effective resolution of incidents and performance issues.
- High .