06 Aug
|
Rakuten Symphony
|
Bengaluru
06 Aug
Rakuten Symphony
Bengaluru
Job Title: MLOps Lead - Enterprise AI Platform
Location: Bangalore (Hybrid)
Responsibilities:
Strategic Platform Architecture:
- Lead the architectural vision, design, and continuous evolution of the Platform, ensuring alignment with business objectives, security standards, and scalability requirements.
- Drive the adoption and integration of open-source MLOps tools (Kubeflow, MLflow, Feast, KServe, Alibi-Detect, Evidently AI, Spark, etc.) into a cohesive, production-ready enterprise solution. • Define platform standards, best practices, and architectural patterns for MLOps development and operations.
Technical Leadership & Implementation Oversight:
- Act as the primary technical authority and lead for the MLOps initiative, guiding both DevOps/Platform and MLOps/Data Science teams through the phased development plan.
- Oversee the implementation of core platform components, ensuring robust integration, performance, and adherence to architectural blueprints.
- Provide expert guidance on Kubernetes-native MLOps practices, distributed computing for ML (Spark, Kubeflow Training Operators), and model serving strategies (KServe).
Enterprise Security, Governance & Multi-Tenancy:
Architect and oversee the implementation of enterprise-grade security features including SSO (Keycloak), secrets management (HashiCorp Vault), and fine-grained access control (Kubernetes RBAC, OPA Gatekeeper) for data and platform resources.
- Design and enforce multi-tenancy models that provide solid isolation, resource governance, and secure data access for internal teams and external customers.
- Ensure the platform meets stringent compliance requirements through comprehensive audit logging, tracing (Fluentd, ELK/OpenSearch, Prometheus/Grafana), and data lineage considerations.
ML Lifecycle & Data Management Expertise:
- Architect and integrate a robust Feature Store (tool like Feast) for consistent feature engineering, management, and serving across training and inference.
- Lead the integration of MLflow for experiment tracking, model versioning, and a centralized model registry.
- Design and implement comprehensive model monitoring solutions, including data drift and model quality detection (Alibi-Detect/Evidently AI), with integrated alerting.
Developer Experience & Customization:
- Champion the developer experience for data scientists, ensuring ease of use, self-service capabilities, and efficient workflows (e.g., automated namespace provisioning, notebook environment management).
- Provide architectural guidance for building a custom, branded UI layer on top of the open source components, enhancing usability and aligning with product offerings.
Collaboration & Mentorship:
- Collaborate extensively with Data Science, DevOps, Security, Product Management, and Business stakeholders to gather requirements, communicate technical vision, and drive platform adoption.
- Mentor and upskill engineering teams in MLOps best practices, cloud-native development, and advanced ML techniques.
Required Skills & Expertise :
- 7+ years of progressive experience in software engineering, data engineering, or MLOps, with at least 5 years in a lead or architect role focused on building and managing production of large-scale ML platforms.
- Expert-level proficiency with Kubernetes and its ecosystem (operators, CRDs, Helm, networking, storage).
- Experience in building/managing ML platform tools such as MLflow , Kubeflow, Airflow, SageMaker, Vertex AI, or Azure Machine Learning.
- Deep hands-on experience with Kubeflow (Pipelines, Notebooks, Training Operators, KServe) in production environments.
- Extensive experience with MLflow for experiment tracking, model registry, and model lifecycle management.
Proven expertise in designing and implementing Feature Stores (e.g., Feast) for both online and offline serving.
- Strong background in distributed data processing technologies like Apache Spark/PySpark, especially on Kubernetes.
- Architectural experience with enterprise security solutions including SSO (Keycloak, OAuth/OIDC), secrets management (HashiCorp Vault), and policy enforcement (Kubernetes RBAC, OPA Gatekeeper).
- Demonstrated ability to implement comprehensive monitoring and observability stacks (Prometheus, Grafana, ELK/OpenSearch, Fluentd, Jaeger) for platform health and ML model performance/drift (Alibi-Detect, Evidently AI).
- Proficiency in Python and experience with major ML/Deep Learning frameworks (TensorFlow, PyTorch, Scikit-learn).
- Experience with cloud-native storage solutions (e.g., MinIO, S3, GCS) and open table formats (Iceberg, Delta Lake).
- Excellent communication, leadership, and interpersonal skills with the ability to influence technical direction and drive complex initiatives across multiple teams.
RAKUTEN SHUGI PRINCIPLES: Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.
- Always improve, always advance. Only be satisfied with complete success - Kaizen.
- Be passionately professional. Take an uncompromising approach to your work and be determined to be the best.
- Hypothesize - Practice - Validate - Shikumika. Use the Rakuten Cycle to success in unknown territory.
- Maximize Customer Satisfaction. The greatest satisfaction for workers in a service industry is to see their customers smile.
- Speed!! Speed!! Speed!! Always be conscious of time. Take charge, set clear goals, and engage your team.
📌 Lead MLOps (Bengaluru)
🏢 Rakuten Symphony
📍 Bengaluru