02 Oct
|
UsefulBI
|
India
We are looking for an AWS MLOps Engineer who can build, deploy, automate, and manage scalable ML workloads in production. The ideal candidate should have strong experience with AWS, SageMaker, MLflow, Docker, CI/CD, and ML model deployment/monitoring.
Key Responsibilities
- Design, build, and manage end-to-end ML/MLOps pipelines for training, validation, deployment, and monitoring.
- Deploy and manage machine learning models in production environments on AWS.
- Build automated workflows for model retraining, model versioning, and deployment.
- Implement model monitoring, data/model drift detection, and performance tracking.
- Develop scalable ML infrastructure using AWS SageMaker and other AWS services.
- Build and manage containerized ML workloads using Docker, Amazon ECR, and ECS/Fargate.
- Orchestrate ML workflows using AWS Step Functions and EventBridge.
- Implement and maintain CI/CD pipelines for ML models and infrastructure.
- Work closely with Data Scientists, Data Engineers, Software Engineers, and DevOps teams to productionize ML solutions.
- Implement infrastructure as code using AWS CDK or CloudFormation.
- Ensure ML systems are scalable, reliable, secure, and maintainable in production.
- Troubleshoot deployment, pipeline, infrastructure, and model-serving issues.
- Establish best practices around MLOps, automation, monitoring, model governance, and release management.
Required Skills
- Robust programming experience in Python.
- Hands-on experience with AWS SageMaker and production ML deployments.
- Strong experience with MLflow for experiment tracking, model management, and/or model deployment.
- Experience with AWS Step Functions for workflow orchestration.
- Hands-on experience with AWS ECS/Fargate and containerized workloads.
- Strong knowledge of Docker and Amazon ECR.
- Experience building and managing CI/CD pipelines.
- Experience with Amazon EventBridge and event-driven architectures.
- Experience with AWS CDK and/or CloudFormation.
- Good understanding of ML lifecycle management, including training, validation, deployment, monitoring, retraining, and rollback.
- Strong understanding of cloud infrastructure, networking, security, and deployment best practices.
Good to Have
- Experience with Kubernetes / Amazon EKS.
- Experience with Terraform.
- Knowledge of GitOps, DevSecOps, and infrastructure automation.
- Experience deploying Generative AI / LLM applications in production.
- Experience with RAG pipelines, model serving, or LLM deployment.
- Familiarity with AWS services such as S3, Lambda, IAM, CloudWatch, and API Gateway.
- Experience working with feature stores, model registries, or model monitoring platforms.
📌 MLOps Engineer (India)
🏢 UsefulBI
📍 India