16 Aug
|
ANSR
|
Bengaluru
ANSR is hiring for one of its clients.
About ANSR MedTech:
Who We Are:
ANSR MedTech Capability Center is a new global innovation hub being established in India for a Fortune 100 Fastest-Growing Company in the MedTech sector. Built in partnership with ANSR, the center draws on ANSR’s proven experience in establishing and scaling high-performance Global Capability Centers (GCCs) for leading global enterprises.
ANSR MedTech center brings together world-class engineering, product, and technology talent to build next-generation healthcare platforms and solutions that power global operations.
Our Vision:
To build a next generation MedTech capability center that powers global healthcare innovation.
We envision:
- High impact innovation hubs shaping global product and technology roadmaps
- Centers that go beyond support functions to drive core engineering and platform development
- Sustainable, scalable ecosystems that nurture world class MedTech talent
- Capability centers that directly influence patient outcomes worldwide
Job Title: Staff ML Ops Engineer
Sub Function: AI Engineering
Location: Bengaluru, India
About the Role:
- The Staff ML Ops Engineer will be a core contributor within the India COE AI engineering practice, responsible for building and operating the platforms, pipelines, and automation that move machine learning models from experimentation into reliable, monitored production.
- This role partners closely with Data Science, AI Engineering, Data Engineering, and Analytics Engineering to ensure models are deployed, versioned, monitored, and governed reproducible and production ready by default. The AI/ML Ops Engineer is expected to execute end to end MLOps work with limited oversight while contributing to shared tooling, reusable infrastructure, and operational standards that raise reliability across the team.
Scope of Responsibility:
Model Deployment & Serving:
- Design, build, and maintain deployment pipelines that promote trained models from development to staging to production supporting batch, real time, and streaming inference patterns
- Package models and their dependencies into reproducible, versioned artifacts (containers, model registry entries) that deploy consistently across environments
- Stand up scalable model serving infrastructure REST/gRPC endpoints, batch scoring jobs, and low latency online inference with appropriate autoscaling and resource management
- Own the model registry and lifecycle versioning, stage transitions, approvals, and archival so every production model is traceable to its training run, data,
and configuration
- Collaborate with Data Science and AI Engineering to operationalize validated models, translating methodology and validation documentation into robust, production grade serving code
Machine Learning Pipelines & CI/CD Automation:
- Build and maintain automated CI/CD pipelines for machine learning covering testing, model validation, packaging, and deployment with quality gates before promotion
- Orchestrate end to end ML workflows (data ingestion, feature generation, training, evaluation, and deployment) using workflow orchestration tooling
- Automate retraining and re-validation cycles triggered by schedule, data drift, or performance degradation with human approval gates where governance requires
- Implement automated testing for ML systems, including unit and integration tests, data validation checks, and model quality thresholds
- Manage feature pipelines and feature stores to keep training and serving consistent and prevent training serving skew
- Codify environments as infrastructure as code so they are reproducible, auditable, and version controlled
Monitoring, Observability & Reliability:
- Instrument deployed models with monitoring for data drift, concept drift, prediction quality, latency, and throughput
- Build alerting and dashboards that surface model degradation, pipeline failures, and SLA breaches to the right owners in time to act
- Define and track operational SLAs/SLOs for model serving systems availability, latency, and error budgets
- Implement logging, tracing, and audit trails for predictions and pipeline runs to support debugging, reproducibility, and compliance
- Lead incident response for production ML systems triage, root cause analysis, remediation, and post incident review
- Partner with Data Science to define the metrics and thresholds that determine when a model should be retrained, rolled back, or retired
ML Platform, Infrastructure & Reproducibility:
- Build and maintain the shared MLOps platform and tooling that raise productivity and the operational standard across the AI practice
- Manage compute infrastructure for training and inference GPU/CPU allocation, cost optimization,
and environment provisioning
- Ensure reproducibility across the ML lifecycle by versioning data, code, models, and configuration so any result can be reconstructed
- Implement security, access control, and governance controls for models, data, and pipelines in line with COE and regulatory standards
- Contribute reusable templates, reference implementations, and self service tooling that reduce time to production for data scientists and AI engineers
- Ensure all deployed systems meet the COE's data product certification and operational readiness standards before release
Required Qualifications:
- Bachelor's or Master's degree in computer science, engineering, or a related technical field or equivalent practical experience
- 8+ years of experience in software engineering, ML engineering, DevOps, or MLOps, with hands on ownership of production systems
- Strong proficiency in Python and solid software engineering fundamentals testing, version control, code review, and modular design
- Hands on experience deploying and operating machine learning models in production model serving, batch scoring, or streaming inference
- Experience with containerization and orchestration (Docker, Kubernetes) and building CI/CD pipelines for ML or software systems
- Experience with workflow orchestration and ML lifecycle tooling (e.g., MLflow, Airflow, Kubeflow, or equivalent) and model registries
- Hands on experience with Databricks for ML workflows, model management, and pipeline orchestration
- Working knowledge of a major cloud platform (AWS, Azure, or GCP) and infrastructure as code practices
- Ability to communicate operational and reliability considerations clearly to technical and non-technical stakeholders
- Experience working in a structured delivery environment with sprint cadences and cross functional collaboration
Preferred Qualifications:
- Experience in a regulated industry (life sciences, healthcare, financial services) where model deployment and monitoring are subject to scrutiny or audit
- Experience building or operating a feature store and eliminating training serving skew
- Familiarity with LLMOps / generative AI operational patterns prompt versioning, evaluation harnesses, and inference cost management
- Experience implementing model governance, lineage, and reproducibility controls at scale
- Experience with observability tooling (Prometheus, Grafana, or equivalent) applied to ML systems
- Experience working in a global capability center (GCC) or center of excellence (COE) setting
📌 Staff Ml Ops Engineer-28545] (Bengaluru)
🏢 ANSR
📍 Bengaluru