MLOps Engineer (Mumbai)

MLOps Engineer (Mumbai)

12 Sep
|
Guddge Tech
|
Mumbai

12 Sep

Guddge Tech

Mumbai

About the Role

We're looking for an MLOps Engineer to support and evolve production machine-learning platform that produces daily risk scores used by operational leaders to intervene before incidents occur. You'll own the reliability, deployment, and monitoring of a production ML pipeline that spans AWS (S3, Glue, SageMaker, Step Functions), and a fully automated Terraform + GitHub Actions delivery pipeline across multi environments.

This is a hands-on infrastructure-and-operations role. The model itself is trained offline; your focus is keeping the scoring pipeline dependable, observable, secure, and easy to change.

What You'll Do

- Operate and improve the end-to-end inference pipeline (SAP HANA --> AWS Glue --> AWS SageMaker --> AWS Step Functions --> SAP HANA), keeping scheduled daily runs reliable.
- Own the infrastructure-as-code: maintain Terraform modules and configuration, manage remote state in Terraform Cloud, and promote changes through dev --> qa --> prd.
- Maintain and extend the CI/CD workflows in GitHub Actions
- Keep the pipeline observable: structured JSON logging, CloudWatch Logs Insights queries, and Step Functions execution monitoring.
- Support model monitoring (drift detection) and data-quality checks; tune alert thresholds and triage failure alerts.
- Troubleshoot production incidents using established runbooks (database connectivity, SageMaker job failures, Glue write-back, scheduler issues) and drive root-cause fixes.
- Enforce security and compliance controls: encryption, least-privilege IAM, secrets handling, and no-PII-in-logs discipline.
- Collaborate with data scientists to deploy retrained models through the SageMaker Model Registry and cross-account promotion process.

Required Qualifications

- 3+ years in MLOps supporting production systems.
- Strong AWS experience,



especially the data/ML services.
- Production Terraform experience with a remote backend and multi-environment promotion.
- Solid Python for data processing and operational scripting.
- Experience operating CI/CD pipelines and diagnosing production failures from logs and metrics.

Technical Skillset Breakdown The percentages reflect the relative weight of each area for day-to-day success in this role. 1. AWS Cloud & ML Services: ~35%

The core of the platform runs on AWS.

- AWS Glue - Spark/PySpark ETL jobs, Glue connections (VPC/JDBC), Data Catalog, crawlers, job bookmarks
- Amazon SageMaker - Processing Jobs, Model Registry, cross-account model packages, execution roles
- AWS Step Functions - orchestration, Amazon States Language (ASL), retry/error handling
- Amazon EventBridge Scheduler - cron scheduling, timezone handling
- Amazon S3 - bucket policies, versioning, server-side encryption, lifecycle
- Amazon CloudWatch - Logs, Logs Insights, monitoring
- AWS Lake Formation - fine-grained catalog access control
- VPC / Security Groups / Prefix Lists network connectivity to on-prem

2. Infrastructure as Code (Terraform): ~20% Nearly all infrastructure is managed as code.

- Terraform module authoring and reuse
- Terraform Cloud (remote state, workspaces, locking)
- Multi-environment configuration (.tfvars, per-env config files)
- Provider configuration (AWS, AWSCC)
- Managing drift, plan/apply safety, state troubleshooting

3.



CI/CD & Automation: ~15% Delivery is fully automated.

- GitHub Actions: workflow authoring, reusable composite actions, environment approvals
- Sequential environment promotion (dev qa prd)
- Git and branch/PR workflows, GitHub CLI / AWS Kiro
- Make for build automation

4. Python & ML Tooling: ~10% Supporting the model code and tests.

- Python 3.8+ for inference/ETL scripts and operational tooling
- pandas, pyarrow, scipy for data processing
- boto3 for AWS automation
- SHAP (model explainability), imbalanced-learn (SMOTE) familiarity to support the model
- Scikit-learn model artifacts (Random Forest, MinMaxScaler) and pickle handling
- pytest + coverage (50% gate), Black, flake8, isort

5. Docker & Containerization: ~10% Containers underpin local testing and reproducible builds.

- Building and running Docker images for development and CI
- Containerized test execution (e.g., PySpark unit tests that don't run natively on Windows)
- Managing dependencies and reproducible environments across local and pipeline runs
- Familiarity with container-based execution in AWS (Glue/SageMaker managed containers)

6. Observability, Reliability & Security: ~10% Keeping it production-grade.

- Structured JSON logging and a typed error hierarchy (config/data/model/inference/downstream)
- Model drift and data-quality monitoring; alert threshold tuning
- Incident triage using runbooks; alerting via SNS/SES dispatcher integration
- Security controls: encryption in transit/at rest, least-privilege IAM, no PII in logs
- Data-quality validation

Working Workplace

- This is a offshore role, you will be working from your base location. No travel required
- Ability to work US Pacific Standard Time Zone (i.e. 8 pm IST to 5 am IST)

📌 MLOps Engineer (Mumbai)
🏢 Guddge Tech
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: mlops engineer (mumbai) / mumbai