01Responsibilities
Collaborate with stakeholders (data scientists, analysts, engineers, business teams) to understand data needs and translate them into actionable technical requirements.
Design, build, and maintain scalable data pipelines using modern tools and cloud platforms (such as AWS, Azure, or GCP).
Develop and maintain data infrastructure for analytics and machine learning, leveraging cloud-native technologies and services.
Improve data quality, lineage, and reliability, and troubleshoot issues across the data lifecycle.
Support the deployment and monitoring of analytics and ML solutions in production, working closely with data science and engineering peers.
Contribute to an Agile team environment, including documentation, code reviews, and iterative delivery.
02Requirements
University degree in Computer Science, Engineering, Mathematics, or a related discipline (or equivalent practical experience).
5+ years of hands-on programming experience, primarily in Python, with a strong focus on building, training, and evaluating machine learning and deep learning models.
Strong experience working with machine learning frameworks and libraries (e.g., PyTorch, TensorFlow, Keras, scikit-learn), including designing and implementing neural network architectures.
Proven experience in computer vision techniques and applications, such as image/video processing, feature extraction, object detection, segmentation, and representation learning.
Experience developing end-to-end ML workflows, from data preparation and feature engineering through model training, validation, and deployment, with an emphasis on model performance, robustness, and maintainability.
Working knowledge of data storage and querying technologies (SQL and/or NoSQL) to support analytical and ML workflows.
Hands-on experience using cloud platforms (such as AWS, Azure, or GCP) for machine learning experimentation, training at scale, and model deployment.
Familiarity with distributed and accelerated computing concepts relevant to ML workloads (e.g., multi-GPU training, distributed training, or large-scale inference).
Experience collaborating closely with data engineers and platform teams to integrate models into production pipelines, while maintaining ownership of model logic and performance.
Exposure to workflow orchestration or experiment management tools (e.g., Airflow, MLflow, or similar) for reproducible ML pipelines is a plus.
Familiarity with version control systems (e.g., Git) and best practices for collaborative ML development.
Strong ability to communicate results, trade-offs, and limitations of ML models clearly to both technical and non-technical stakeholders.
Nice to Have
Experience with data governance practices (priv .