Job Title: ML Ops Engineer
Responsibilities:
1. Model Deployment & Integration:
- Design, develop, and manage automated pipelines for deploying machine learning models into production.
- Ensure smooth integration between model development, data, and application teams.
- Implement model versioning and rollback strategies to facilitate easy model updates and troubleshooting.
2. Infrastructure Automation:
- Build and maintain scalable infrastructure using tools like Kubernetes, Docker, and cloud platforms (AWS, Azure, GCP).
- Automate the deployment process and manage model serving environments.
- Design and optimize cloud-native solutions to ensure scalability and performance under heavy workloads.
3. Monitoring and Maintenance:
- Continuously monitor the performance and health of deployed models in production environments.
- Implement real-time logging, alerting, and monitoring systems to ensure models effectiveness over time.
- Detect, troubleshoot, and resolve issues such as model drift, degradation, and inefficiencies.
4. Collaboration with Data Scientists & DevOps:
- Work closely with data scientists to ensure that models are production-ready and meet system requirements.
- Collaborate with DevOps teams to integrate MLOps tools and practices into the CI/CD pipeline.
- Optimize model performance by coordinating with various teams to manage the lifecycle of machine learning models.
5. Model Retraining & Continuous Improvement:
- Automate and manage model retraining processes based on incoming current data or changing business needs.
- Create frameworks for evaluating and improving model accuracy, efficiency, and robustness.
6. Security & Compliance:
- Ensure the security of machine learning systems, including data protection, model access control, and sensitive data handling.
- Ensure compliance with relevant regulatory requirements related to data privacy and security.
7. Performance Optimization:
- Work on optimizing models and system performance for faster inference