12 Sep
|
Guddge Tech
|
Mumbai
12 Sep
Guddge Tech
Mumbai
About the Role
We're looking for an MLOps Engineer to support and evolve production machine-learning platform that produces daily risk scores used by operational leaders to intervene before incidents occur. You'll own the reliability, deployment, and monitoring of a production ML pipeline that spans AWS (S3, Glue, SageMaker, Step Functions), and a fully automated Terraform + GitHub Actions delivery pipeline across multi environments.
This is a hands-on infrastructure-and-operations role. The model itself is trained offline; your focus is keeping the scoring pipeline dependable, observable, secure, and easy to change.
What You'll Do
- Operate and improve the end-to-end inference pipeline (SAP HANA --> AWS Glue --> AWS SageMaker --> AWS Step Functions --> SAP HANA), keeping scheduled daily runs reliable.
- Own the infrastructure-as-code: maintain Terraform modules and configuration, manage remote state in Terraform Cloud, and promote changes through dev --> qa --> prd.
- Maintain and extend the CI/CD workflows in GitHub Actions
- Keep the pipeline observable: structured JSON logging, CloudWatch Logs Insights queries, and Step Functions execution monitoring.
- Support model monitoring (drift detection) and data-quality checks; tune alert thresholds and triage failure alerts.
- Troubleshoot production incidents using established runbooks (database connectivity, SageMaker job failures, Glue write-back, scheduler issues) and drive root-cause fixes.
- Enforce security and compliance controls: encryption, least-privilege IAM, secrets handling, and no-PII-in-logs discipline.
- Collaborate with data scientists to deploy retrained models through the SageMaker Model Registry and cross-account promotion process.
Required Qualifications
- 3+ years in MLOps supporting production systems.
- Strong AWS experience,
especially the data/ML services.
- Production Terraform experience with a remote backend and multi-environment promotion.
- Solid Python for data processing and operational scripting.
- Experience operating CI/CD pipelines and diagnosing production failures from logs and metrics.
Technical Skillset Breakdown The percentages reflect the relative weight of each area for day-to-day success in this role. 1. AWS Cloud & ML Services: ~35%
The core of the platform runs on AWS.
- AWS Glue - Spark/PySpark ETL jobs, Glue connections (VPC/JDBC), Data Catalog, crawlers, job bookmarks
- Amazon SageMaker - Processing Jobs, Model Registry, cross-account model packages, execution roles
- AWS Step Functions - orchestration, Amazon States Language (ASL), retry/error handling
- Amazon EventBridge Scheduler - cron scheduling, timezone handling
- Amazon S3 - bucket policies, versioning, server-side encryption, lifecycle
- Amazon CloudWatch - Logs, Logs Insights, monitoring
- AWS Lake Formation - fine-grained catalog access control
- VPC / Security Groups / Prefix Lists network connectivity to on-prem
2. Infrastructure as Code (Terraform): ~20% Nearly all infrastructure is managed as code.
- Terraform module authoring and reuse
- Terraform Cloud (remote state, workspaces, locking)
- Multi-environment configuration (.tfvars, per-env config files)
- Provider configuration (AWS, AWSCC)
- Managing drift, plan/apply safety, state troubleshooting
3.
CI/CD & Automation: ~15% Delivery is fully automated.
- GitHub Actions: workflow authoring, reusable composite actions, environment approvals
- Sequential environment promotion (dev qa prd)
- Git and branch/PR workflows, GitHub CLI / AWS Kiro
- Make for build automation
4. Python & ML Tooling: ~10% Supporting the model code and tests.
- Python 3.8+ for inference/ETL scripts and operational tooling
- pandas, pyarrow, scipy for data processing
- boto3 for AWS automation
- SHAP (model explainability), imbalanced-learn (SMOTE) familiarity to support the model
- Scikit-learn model artifacts (Random Forest, MinMaxScaler) and pickle handling
- pytest + coverage (50% gate), Black, flake8, isort
5. Docker & Containerization: ~10% Containers underpin local testing and reproducible builds.
- Building and running Docker images for development and CI
- Containerized test execution (e.g., PySpark unit tests that don't run natively on Windows)
- Managing dependencies and reproducible environments across local and pipeline runs
- Familiarity with container-based execution in AWS (Glue/SageMaker managed containers)
6. Observability, Reliability & Security: ~10% Keeping it production-grade.
- Structured JSON logging and a typed error hierarchy (config/data/model/inference/downstream)
- Model drift and data-quality monitoring; alert threshold tuning
- Incident triage using runbooks; alerting via SNS/SES dispatcher integration
- Security controls: encryption in transit/at rest, least-privilege IAM, no PII in logs
- Data-quality validation
Working Workplace
- This is a offshore role, you will be working from your base location. No travel required
- Ability to work US Pacific Standard Time Zone (i.e. 8 pm IST to 5 am IST)
📌 MLOps Engineer (Mumbai)
🏢 Guddge Tech
📍 Mumbai