01Key Responsibilities
Architect and implement the MLOps strategy for the programme, ensuring alignment with the project proposal and delivery roadmap.
Design and own enterprise-grade ML/LLM pipelines covering model training, validation, deployment, versioning, monitoring, and CI/CD automation.
Build container-oriented ML platforms (EKS-first) while evaluating alternative orchestration tools with similar capabilities (Kubeflow, SageMaker, MLflow, Airflow, etc.).
Implement hybrid MLOps + LLMOps workflows, including prompt/version governance, evaluation frameworks, and monitoring for LLM-based systems.
Serve as a technical authority across multiple internal and customer projects, contributing architectural patterns, best practices, and reusable frameworks.
Enable observability, monitoring, drift detection, lineage tracking, and auditability across ML/LLM systems.
Define and implement standards for model deployment, monitoring, governance, and automation to ensure production-grade reliability and scalability.
Collaborate with cross-functional teams data engineering, platform, DevOps, and client stakeholders to deliver production-ready ML solutions.
Ensure all solutions adhere to security, governance, and compliance expectations, particularly around handling cloud services, Kubernetes workloads, and MLOps tools.
Conduct architecture reviews, troubleshoot complex ML system issues, and guide teams through implementation across cloud-native ML platforms.
Mentor engineers and provide guidance on modern MLOps tools, platform capabilities, and best practices.
Basic Qualifications (BQs):
Expierence working in ML/AI engineering or MLOps roles with strong architecture exposure.
Strong experience in leading enterprise grade MLOps strategy and its execution.
Proven experience in implementing the adoption of enterprise-grade MLOps platforms with client data science teams.
Proven leadership in defining and executing enterprise MLOps strategy.
Demonstrated success in driving the adoption of enterprise-grade MLOps platforms with client data science teams.
Strong expertise in AWS cloud-native ML stack, including: SageMaker(primary), EKS, Lambda, API Gateway, CI/CD (CodeBuild/CodePipeline or equivalent).
Hands-on experience with at least one major MLOps toolset and awareness of alternatives: MLflow, Kubeflow, SageMaker Pipelines, Airflow, BentoML, KServe, Seldon.
Deep understanding of model lifecycle management (feature engineering->training registry deployment monitoring).
Experience implementing or supporting LLMOps pipelines, including: prompt versioning, evaluation metrics, automation frameworks.
Deep understanding of ML lifecycle: data ingestion, feature engineering, training .