- Robust understanding of the end-to-end machine learning lifecycle (data preparation, training, validation, deployment, and monitoring).
- Experience building CI/CD and continuous training pipelines for ML workflows.
- Hands-on expertise with Docker, Kubernetes, MLflow, and Kubeflow.
- Monitoring and observability of ML systems, including model drift and performance tracking.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Knowledge of security, governance, and compliance for production ML systems.
- Programming skills in Python, Java, or .NET, with frameworks such as TensorFlow and PyTorch.