Job Description: MLOps & AI Platform Engineer:
Job Title: MLOps & AI Platform Engineer
Experience: 3-11 Years
Location:Riyadh - Onsite
Employment Type: Full-Time
Job Overview:
Key Responsibilities:
- Design, build, and maintain enterprise-grade MLOps platforms and AI infrastructure.
- Develop and automate end-to-end machine learning pipelines for training, validation, deployment, and monitoring.
- Implement model versioning, experiment tracking, and model registry solutions.
- Build scalable CI/CD pipelines for AI/ML workloads.
- Deploy and manage machine learning workloads on Kubernetes-based environments.
- Collaborate with Data Scientists, AI Engineers, Data Engineers, and DevOps teams to operationalize ML solutions.
- Implement Infrastructure as Code (IaC) for cloud-native AI platforms.
- Monitor platform health, model performance, and infrastructure availability.
- Ensure platform security, scalability, reliability, and operational excellence.
- Troubleshoot production issues and continuously optimize platform performance.
Required Technical Skills:
MLOps Platforms:
- Hands-on experience with Kubeflow or Vertex AI Pipelines or SageMaker Pipelines.
- Strong experience with MLflow for experiment tracking, model registry, and lifecycle management.
- Experience orchestrating machine learning workflows using Apache Airflow.
Containerization & Orchestration:
- Solid expertise in Kubernetes (GKE or AKS or EKS).
- Experience deploying and managing containerized AI/ML workloads in cloud environments.
Infrastructure Automation:
- Hands-on experience with Terraform for Infrastructure as Code (IaC).
- Experience automating infrastructure provisioning and cloud resource management.
CI/CD & DevOps:
- Experience with GitHub Actions for CI/CD automation.
- Knowledge of DevOps best practices, Git workflows, and automated deployments.
Monitoring & Observability:
- Knowledge of logging, alerting, and performance monitoring for AI platforms.