12 Sep
|
Guddge Tech
|
Mumbai
12 Sep
Guddge Tech
Mumbai
About the Role :We're looking for an MLOps Engineer to support and evolve a production machine-learning platform that produces daily risk scores used by operational leaders to intervene before incidents occur. You'll own the reliability, deployment, and monitoring of a production ML pipeline that spans AWS (S3, Glue, SageMaker, Step Functions), and a fully automated Terraform + GitHub Actions delivery pipeline across multi environments.This is a hands-on infrastructure-and-operations role. The model itself is trained offline; your focus is keeping the scoring pipeline dependable, observable, secure, and easy to change. What You'll Do :- Operate and improve the end-to-end inference pipeline (SAP HANA - AWS Glue - AWS SageMaker - AWS Step Functions - SAP HANA), keeping scheduled daily runs reliable.- Own the infrastructure-as-code: maintain Terraform modules and configuration, manage remote state in Terraform Cloud, and promote changes through dev - qa - prd.- Maintain and extend the CI/CD workflows in GitHub Actions.- Keep the pipeline observable: structured JSON logging, CloudWatch Logs Insights queries, and Step Functions execution monitoring.- Support model monitoring (drift detection) and data-quality checks; tune alert thresholds and triage failure alerts.- Troubleshoot production incidents using established runbooks (database connectivity, SageMaker job failures, Glue write-back, scheduler issues) and drive root-cause fixes.- Enforce security and compliance controls: encryption, least-privilege IAM, secrets handling, and no-PII-in-logs discipline.- Collaborate with data scientists to deploy retrained models through the SageMaker Model Registry and cross-account promotion process.Required Qualifications :- 3+ years in MLOps supporting production systems.- Strong AWS experience,
especially the data/ML services.- Production Terraform experience with a remote backend and multi-environment promotion.- Solid Python for data processing and operational scripting.- Experience operating CI/CD pipelines and diagnosing production failures from logs and metrics.Technical Skillset Breakdown :The percentages reflect the relative weight of each area for day-to-day success in this role.1. AWS Cloud & ML Services: ~35% :The core of the platform runs on AWS.- AWS Glue - Spark/PySpark ETL jobs, Glue connections (VPC/JDBC), Data Catalog, crawlers, job bookmarks.- Amazon SageMaker - Processing Jobs, Model Registry, cross-account model packages, execution roles.- AWS Step Functions - orchestration, Amazon States Language (ASL), retry/error handling.- Amazon EventBridge Scheduler - cron scheduling, timezone handling.- Amazon S3 - bucket policies, versioning, server-side encryption, lifecycle.- Amazon CloudWatch - Logs, Logs Insights, monitoring.- AWS Lake Formation - fine-grained catalog access control.- VPC / Security Groups / Prefix Lists - network connectivity to on-prem.2. Infrastructure as Code (Terraform): ~20% :Nearly all infrastructure is managed as code.- Terraform module authoring and reuse.- Terraform Cloud (remote state, workspaces, locking).- Multi-environment configuration (.tfvars, per-env config files).- Provider configuration (AWS, AWSCC).- Managing drift, plan/apply safety, state troubleshooting.3.
CI/CD & Automation: ~15% : Delivery is fully automated.- GitHub Actions: workflow authoring, reusable composite actions, environment approvals.- Sequential environment promotion (dev - qa - prd).- Git and branch/PR workflows, GitHub CLI / AWS Kiro.- Make for build automation.4. Python & ML Tooling: ~10% :Supporting the model code and tests.- Python 3.8+ for inference/ETL scripts and operational tooling.- pandas, pyarrow, scipy for data processing.- boto3 for AWS automation.- SHAP (model explainability), imbalanced-learn (SMOTE) - familiarity to support the model.- Scikit-learn model artifacts (Random Forest, MinMaxScaler) and pickle handling.- pytest + coverage (50% gate), Black, flake8, isort.5. Docker & Containerization: ~10% :Containers underpin local testing and reproducible builds.- Building and running Docker images for development and CI.- Containerized test execution (e.g., PySpark unit tests that don't run natively on Windows).- Managing dependencies and reproducible environments across local and pipeline runs.- Familiarity with container-based execution in AWS (Glue/SageMaker managed containers).6. Observability, Reliability & Security: ~10% :Keeping it production-grade.- Structured JSON logging and a typed error hierarchy Model drift and data-quality monitoring; alert threshold tuning.- Incident triage using runbooks; alerting via SNS/SES dispatcher integration.- Security controls: encryption in transit/at rest, least-privilege IAM, no PII in logs.- Data-quality validation.Working Setting :- This is an offshore role, you will be working from your base location. No travel required.- Ability to work US Pacific Standard Time Zone (i.e. 8 pm IST to 5 am IST). (ref:hirist.tech)
📌 Guddge - MLOps Engineer (Mumbai)
🏢 Guddge Tech
📍 Mumbai