02 Oct
|
Affine Analytics
|
Bengaluru
02 Oct
Affine Analytics
Bengaluru
Job Summary
You own the path to production and the production system itself - turning modeling work into reliable, scalable, observable services that run at Expedias traffic volumes. You partner closely with ML Specialists/Data Scientists, but the accountability for latency, uptime, cost, retraining pipelines, and CI/CD for models is yours.
This role suits someone who is as comfortable in distributed systems and MLOps as in ML itself - think software engineer who specializes in ML systems rather than data scientist who can code.
Responsibilities
- Design, build, and operate production ML pipelines: feature engineering, training pipelines, model serving, and monitoring - at Expedias scale (millions of requests/day).
- Build and maintain low-latency, high-availability model serving infrastructure (real-time inference for search/personalization; batch scoring for pricing/forecasting).
- Own CI/CD for ML: automated retraining, model versioning, shadow deployments, canary releases, and rollback strategies.
- Build robust feature pipelines (batch and streaming) using Spark and Kafka, and maintain feature stores for training/serving consistency.
- Instrument models and pipelines with observability: drift detection, data quality checks, latency/throughput SLAs, and alerting.
- Collaborate with ML Specialists to productionize research prototypes - translating notebook code into tested, maintainable, scalable services.
- Optimize training and inference cost/performance (GPU utilization, batching, caching, model compression/quantization where relevant).
- Contribute to platform-level decisions: build vs. buy for MLOps tooling,
standards for reproducibility, and shared infrastructure across ML teams.
- Participate in on-call rotation for production ML services; drive postmortems and reliability improvements.
- Mentor mid-level ML/software engineers and lead design reviews for ML platform and serving infrastructure components.
Must-Have Qualifications
- 6-9 years of software engineering experience, with 4+ years specifically building and operating production ML systems.
- Robust software engineering fundamentals: Java, Scala, or Python at a production quality bar (testing, code review, design patterns), plus working knowledge of the other two.
- Deep experience with cloud infrastructure (AWS preferred) - EC2, EKS/Kubernetes, S3, Lambda, IAM - and infrastructure-as-code (Terraform or similar).
- Event-Driven Systems: Experience building and operating streaming/event-driven pipelines (Kafka) and distributed batch processing (Spark), including microservices that consume and produce events at scale.
- Hands-on experience with model serving frameworks (e.g., TorchServe, TensorFlow Serving, Triton, or custom microservices) and CI/CD for ML (MLflow, SageMaker Pipelines, Airflow, or similar).
- Testing Data Quality:
Experience with data validation and testing practices for ML pipelines - unit/integration tests for data and models, and contract tests between training and serving to prevent train/serve skew.
- Security Compliance: Experience handling PII and other sensitive data securely within ML pipelines; familiarity with data governance, access controls (IAM policies), and encryption at rest/in transit.
- Solid understanding of ML fundamentals (enough to have real technical conversations with data scientists) even if you dont build models from scratch day-to-day.
- Technical Leadership: Track record of mentoring engineers and leading design/architecture reviews for distributed or ML-platform systems.
- Track record of owning production systems: SLAs, on-call, incident response, capacity planning.
Nice-to-Have
- Experience with feature stores (Feast, Tecton, or internal equivalents) and real-time feature computation.
- Familiarity with AWS SageMaker, Databricks, or Kubeflow for orchestration.
- Experience with GPU infrastructure and inference optimization (quantization, distillation, batching strategies).
- Background in search/ranking, personalization, pricing, or fraud systems at marketplace scale.
- Experience with A/B testing infrastructure for ML models (shadow traffic, interleaving, canary analysis).
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Consultant - Machine Learning Engineer (Bengaluru)
🏢 Affine Analytics
📍 Bengaluru