01Responsibilities
Lead end-to-end design and implementation of AI Ops solutions from concept through production with an emphasis on responsible AI practicesContribute to the improvement of our SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environmentsAutomate Infrastructure & Operations: Develop Infrastructure as Code using Terraform & GitHub Actions while adhering to best practicesDefine and own solution architecture for RAG pipelines, agentic workflows, tool calling, orchestration, and conversational context managementDesign, develop, and deploy AI-powered solutions to address complex business challenges across enterprise scale using RAG based solutions Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineersDesign, code, test, and operate software using Python & Node.jsLeverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous operational efficiencyComply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do soRequired Qualifications:
Bachelor's degree in information systems, Computer Science, Engineering, or related field or equivalent certification5+ years of overall software engineering and Site Reliability Engineering (SRE) experience in a public cloud environment (GCP, AWS, Azure) 3+ years of demonstrated hands-on experience with Python and Terraform based development 2+ years delivering AI/ML or Generative AI solutions in production 2+ years of experience managing Kubernetes environments (EKS, AKS, GKE, or self-hosted) Available to work rotating 24x7 primary and secondary on-call shifts
Preferred Qualifications:
Hands-on experience with cloud infrastructure automation, observability tools, and SRE best practices for AI workloads Experience leading globally distributed technical teams Experience with vector search and enterprise search solutions CI/CD (GitHub Actions preferred) & Infrastructure as Code experience with Terraform Proven experience building Retrieval Augmented Generation (RAG) pipelines, Agentic AI, or multi-step AI workflows Proven experience building evaluation and monitoring frameworks for AI quality Proven effective communication skills with ability to explain complex technical concepts to diverse stakeholders
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission. .