27 Sep
|
Affine Analytics
|
Bengaluru
27 Sep
Affine Analytics
Bengaluru
- Design, build, and operate production ML pipelines: feature engineering, training pipelines, model serving, and monitoring at Expedia's scale (millions of requests/day).
- Build and maintain low-latency, high-availability model serving infrastructure (real-time inference for search/personalization; batch scoring for pricing/forecasting).
- Own CI/CD for ML: automated retraining, model versioning, shadow deployments, canary releases, and rollback strategies.
- Build robust feature pipelines (batch and streaming) using Spark and Kafka, and maintain feature stores for training/serving consistency.
- Instrument models and pipelines with observability: drift detection, data quality checks, latency/throughput SLAs, and alerting.
- Collaborate with ML Specialists to productionize research prototypes — translating notebook code into tested, maintainable, scalable services.
- Optimize training and inference cost/performance (GPU utilization, batching, caching, model compression/quantization where relevant).
- Contribute to platform-level decisions: build vs. buy for MLOps tooling, standards for reproducibility, and shared infrastructure across ML teams.
- Participate in on-call rotation for production ML services; drive postmortems and reliability improvements.
- Mentor mid-level ML/software engineers and lead design reviews for ML platform and serving infrastructure components.
Must-Have Qualifications
- 6–9 years of software engineering experience, with 4+ years specifically building and operating production ML systems.
- Robust software engineering fundamentals: Java, Scala, or Python at a production quality bar (testing, code review, design patterns), plus working knowledge of the other two.
- Deep experience with cloud infrastructure (AWS preferred) — EC2, EKS/Kubernetes, S3, Lambda, IAM — and infrastructure-as-code (Terraform or similar).
- Event-Driven Systems: Experience building and operating streaming/event-driven pipelines (Kafka) and distributed batch processing (Spark), including microservices that consume and produce events at scale.
- Hands-on experience with model serving frameworks (e.g., TorchServe, TensorFlow Serving, Triton, or custom microservices) and CI/CD for ML (MLflow, SageMaker Pipelines, Airflow, or similar).
- Testing & Data Quality: Experience with data validation and testing practices for ML pipelines — unit/integration tests for data and models, and contract tests between training and serving to prevent train/serve skew.
- Security & Compliance: Experience handling PII and other sensitive data securely within ML pipelines; familiarity with data governance, access controls (IAM policies), and encryption at rest/in transit.
- Solid understanding of ML fundamentals (enough to have real technical conversations with data scientists) even if you don't build models from scratch day-to-day.
- Technical Leadership: Track record of mentoring engineers and leading design/architecture reviews for distributed or ML-platform systems.
- Track record of owning production systems: SLAs, on-call, incident response, capacity planning.
Nice-to-Have
- Experience with feature stores (Feast, Tecton, or internal equivalents) and real-time feature computation.
- Familiarity with AWS SageMaker, Databricks, or Kubeflow for orchestration.
- Experience with GPU infrastructure and inference optimization (quantization, distillation, batching strategies).
- Background in search/ranking, personalization, pricing, or fraud systems at marketplace scale.
- Experience with A/B testing infrastructure for ML models (shadow traffic, interleaving, canary analysis).
SKILLS AND EXPERTISE
Languages
Java/Scala/Python (production services), SQL
Orchestration/Serving
Kubernetes (EKS), Airflow, internal EGAP platform services, TorchServe/TensorFlow Serving or gRPC-based custom serving
Data Platform
Kafka, Spark, S3 data lake, Presto/Trino, Databricks
MLOps
MLflow, SageMaker Pipelines, Terraform, Docker, CI/CD via Jenkins/GitHub Actions
Cloud
AWS (EC2, EKS, Lambda, S3, IAM)
Observability
CloudWatch/Datadog-style metrics, custom drift/data-quality monitors
Collaboration
Jira/Confluence (Atlassian), Git-based workflows
📌 Machine Learning Engineer (Bengaluru)
🏢 Affine Analytics
📍 Bengaluru