30 Sep
|
CareerVitaBLR
|
Bengaluru
30 Sep
CareerVitaBLR
Bengaluru
You own the path to production and the production system itself turning modeling work into reliable, scalable, observable services that run at client's traffic volumes.
You partner closely with ML Specialists/Data Scientists, but the accountability for latency, uptime, cost, retraining pipelines, and CI/CD for models is yours.
This role suits someone who is as comfortable in distributed systems and MLOps as in ML itself think software engineer who specializes in ML systems rather than data scientist who can code.
Design, build, and operate production ML pipelines: feature engineering, training pipelines, model serving, and monitoring at client's scale (millions of requests/day).
Build and maintain low-latency, high-availability model serving infrastructure (real-time inference for search /personalization; batch scoring for pricing/forecasting).
Own CI/CD for ML: automated retraining, model versioning, shadow deployments, canary releases, and rollback strategies.
Build robust feature pipelines (batch and streaming) using Spark and Kafka, and maintain feature stores for training /serving consistency.
Instrument models and pipelines with observability: drift detection, data quality checks, latency/throughput SLAs, and alerting.
Collaborate with ML Specialists to productionize research prototypes translating notebook code into tested, maintainable, scalable services.
Optimize training and inference cost/performance (GPU utilization, batching, caching, model compression /quantization where relevant).
Contribute to platform-level decisions: build vs. buy for MLOps tooling, standards for reproducibility, and shared infrastructure across ML teams.
Participate in on-call rotation for production ML services; drive postmortems and reliability improvements.
Mentor mid-level ML/software engineers and lead design reviews for ML platform and serving infrastructure components.
Must-Have Qualifications
6 to 9 years of software engineering experience, with 4+ years specifically building and operating production ML systems.
Robust software engineering fundamentals: Java, Scala, or Python at a production quality bar (testing, code review, design patterns), plus working knowledge of the other two.
Deep experience with cloud infrastructure (AWS preferred) EC2, EKS/Kubernetes, S3, Lambda, IAM and infrastructure-as-code (Terraform or similar).
Event-Driven Systems: Experience building and operating streaming/event-driven pipelines (Kafka) and distributed batch processing (Spark), including microservices that consume and produce events at scale.
Hands-on experience with model serving frameworks (e.g., TorchServe, TensorFlow Serving, Triton, or custom microservices) and CI/CD for ML (MLflow, SageMaker Pipelines, Airflow, or similar).
Testing & Data Quality: Experience with data validation and testing practices for ML pipelines unit/integration tests for data and models, and contract tests between training and serving to prevent train/serve skew.
Security & Compliance:
Experience handling PII and other sensitive data securely within ML pipelines; familiarity with data governance, access controls (IAM policies), and encryption at rest/in transit.
Solid understanding of ML fundamentals (enough to have real technical conversations with data scientists) even if you don't build models from scratch day-to-day.
Technical Leadership: Track record of mentoring engineers and leading design/architecture reviews for distributed or ML-platform systems.
Track record of owning production systems: SLAs, on-call, incident response, capacity planning.
Nice-to-Have
Experience with feature stores (Feast, Tecton, or internal equivalents) and real-time feature computation.
Familiarity with AWS SageMaker, Databricks, or Kubeflow for orchestration.
Experience with GPU infrastructure and inference optimization (quantization, distillation, batching strategies).
Background in search/ranking, personalization, pricing, or fraud systems at marketplace scale.
Experience with A/B testing infrastructure for ML models (shadow traffic, interleaving, canary analysis).
EXPERTISE AND QUALIFICATIONS
Languages
Java/Scala/Python (production services), SQL
Orchestration/Serving
Kubernetes (EKS), Airflow, internal EGAP platform services, TorchServe/TensorFlow Serving or gRPC-based custom serving
Data Platform
Kafka, Spark, S3 data lake, Presto/Trino, Databricks
MLOps
MLflow, SageMaker Pipelines, Terraform, Docker, CI/CD via Jenkins/GitHub Actions
Cloud
AWS (EC2, EKS, Lambda, S3, IAM)
Observability
CloudWatch/Datadog-style metrics, custom drift/data-quality monitors
Collaboration
Jira/Confluence (Atlassian), Git-based workflows
📌 Consultant Machine Learning Engineer (Bengaluru)
🏢 CareerVitaBLR
📍 Bengaluru