01Responsibilities
Architectural Pattern Definition: Define platform and model development patterns for ML models, LLMs, and AI agents in collaboration with the Data and AI Architecture team.
Align to the PepsiCo AI Observability patterns to enable robust model monitoring, explainability, and drift detection.
Cross-Functional Collaboration: Engage with cross-functional teams (data science, engineering, operations, etc.) to standardize AI architecture practices and patterns across the enterprise.
Enterprise Alignment: Work closely with the Enterprise Architecture (EA) group to align AI/ML architecture patterns with PepsiCo's broader IT strategy and standards.
Scalability & Reliability: Ensure AI solutions are scalable, reliable, secure, and high-performing across cloud environments (Azure AWS and GCP), following cloud-native architecture best practices.
Innovation & Strategy: Stay up to date with emerging AI/ML trends and technologies and provide strategic technical direction for future AI initiatives and capabilities.
Ensure ecosystem availability and performance in production environments, Pro-actively preventing P1, P2, potential P3s.
Engage & influence product and engineering teams during the design and development phases to embed reliability and operability into new services defining & enforce events, logging, monitoring, and observability standards across applications.
Work closely with customer-facing support teams to empower them with SRE insights and tooling.
Observe, diagnose & improve the end-2-end ecosystem performance of the Modern architected application portfolio i.e. technical understanding of interactions" of a full stack application alongside with peer SRE team member.
Continuously optimize the L2/support operations work via AI and Agentic flows.
Be the key architect for the AI flavor and use cases in SRE orchestration platform design with inputs from Production Operations, Business usage & Product and engineering teams.
Actively engage and drive AI Ops adoption across teams
Qualifications:
Education: Bachelors' / master's in computer science, or a related field.
Experience: 12+ years of experience in developing proactive, diagnosis solutions with a focus on designing and deploying enterprise-scale resiliency/ AI solutions.
Cloud Expertise: Strong expertise in Azure and AWS AI/ML services and cloud-native architectures for machine learning.
Advanced AI Technologies: Hands-on experience with LLMs, AI agents, Machine Learning Operations (ML Ops) pipelines
Programming & Frameworks: Strong programming skills in Python and experience with AI/ML frameworks such as TensorFlow, PyTorch, or similar libraries.
Collaboration: Ability to collaborate effectively with data engineers, data scientists/AI scientists, and IT leadership to drive AI initiatives forward.
Governance & Compliance: Solid understanding of enterprise AI governance, security best practices, and compliance requirements (data privacy, model ethics, etc.).
Exposure to popular AI frameworks like langchain, langgraph and ethical AI practices.
Strong communication skills with the ability to explain complex technical concepts to non-technical stakeholders. .