Role Description We are looking for a Machine Learning Engineer / MLOps Engineer with solid expertise in Python, machine learning model automation, and MLOps. The ideal candidate will have hands-on experience building and maintaining automated machine learning pipelines, managing model artifacts and versions, and deploying ML workflows on cloud platforms such as GCP or Azure ML. The role involves developing scalable and reliable ML infrastructure that supports the complete machine learning lifecycle—from data preparation and model training to validation, versioning, deployment, monitoring, and automated retraining.
The successful candidate will work closely with Data Scientists, ML Engineers, Data Engineers, Software Engineers, and Product teams to productionize machine learning models and establish robust MLOps practices.
Key Responsibilities 1.
Python Development
Develop production-quality machine learning and MLOps solutions using Python. Build reusable Python libraries, utilities, and automation frameworks for ML workflows. Develop scripts and services for model training, validation, deployment, monitoring, and retraining. Write clean, maintainable, testable, and well-documented Python code. Implement appropriate logging, exception handling, configuration management, and testing.
Optimize
Python-based ML workflows for scalability and reliability. 2.
Machine Learning Modeling Automation
Automate the end-to-end machine learning lifecycle.
Build automated workflows for: Data preparation and validation Feature engineering Model training Hyperparameter tuning Model evaluation and validation Model packaging Model registration Deployment Monitoring Model retraining Develop reusable automation components that can support multiple ML models and use cases.
Implement automated model validation and quality checks before deployment. Establish standardized processes for moving models from experimentation to production.
Reduce manual intervention in model development and deployment workflows.
- MLOps & Model Training Pipelines Design, implement, and maintain production-grade ML training pipelines. Build pipelines that can execute repeatable model training jobs using configurable parameters. Automate data ingestion, preprocessing, feature generation, training, evaluation, and model registration. Implement pipeline dependencies, failure handling, retries, logging, and monitoring. Ensure ML pipelines are reproducible across development, testing, and production environments. Integrate automated testing and validation into ML workflows. Collaborate with Data Scientists to convert experimental notebooks and prototypes into production-ready pipelines. 4.
Model Artifact Versioning
Implement robust model artifact versioning and lifecycle management. Version trained models, configuration files, preprocessing logic, feature definitions, and other relevant ML artifacts.
Maintain traceability between: Model version Training dataset/version Code version Configuration Features Training parameters Evaluation results Support reproducibility of historical model versions. Establish processes for promoting models through development, staging, and production environments.
Maintain model registries and ensure appropriate metadata is captured for each model version. Support rollback to previously validated model versions when required.
- GCP / Azure ML Build and operate machine learning workflows on Google Cloud Platform (GCP) and/or Azure Machine Learning. Develop cloud-based training and deployment pipelines. Configure and manage scalable compute resources for ML workloads. Integrate cloud storage, model registries, compute, monitoring, and other ML services into automated workflows. Implement secure and reliable deployment processes for machine learning models. Optimize cloud-based ML workloads for performance, scalability, and cost. Troubleshoot failures across cloud-based training and deployment pipelines. 6.
Model
Deployment & Productionization Work with Data Scientists to productionize machine learning models. Develop standardized processes for packaging and deploying models. Automate model promotion from development to production. Implement appropriate deployment validation and health checks. Support batch and/or real-time inference workflows. Monitor deployed models for availability, performance, and data/model quality. Collaborate with engineering teams to resolve production issues.
- Monitoring & Model Lifecycle Management Establish monitoring processes for ML pipelines and deployed models. Monitor model performance, data quality, pipeline failures, and operational metrics. Implement s for failed training jobs, degraded model performance, or data quality issues. Maintain model lineage and lifecycle information. Define processes for model retirement, replacement, rollback, and re-deployment. Support governance and auditability requirements for production ML systems.
Skills Google Cloud Platform, Azure Machine Learning, MLOps, Python
📌 Lead I - ML Engineering (Bengaluru)
🏢 UST
📍 Bengaluru