02 Aug
|
Rakuten Symphony
|
Bengaluru
02 Aug
Rakuten Symphony
Bengaluru
Job Title: Principal Engineer – AI ML JOB PURPOSE: This section should summarise the purpose of the role, its’ level of business responsibility and strategic input to the business As Principal Engineer in the AIML Framework team, you are the senior technical authority for Rakuten Symphony's complete AIML stack, spanning applied AI/ML delivery and GenAI engineering through to the platform infrastructure, quality systems, and governance frameworks that make them trustworthy at enterprise scale. You will envision, architect, and build the centralised AI/ML platform; drive production GenAI and ML engineering across product teams; and define how AI systems are evaluated, governed, secured, and made explainable. Your software engineering depth is foundational.
It is what enables you to translate architectural intent into working, production grade systems.
PRINCIPLE RESPONSIBILITIES: 1. AI/ML Engineering & GenAI Delivery • Design, build, and operate production AI/ML systems spanning the full model lifecycle: data pipelines,
model training and fine tuning, GenAI/LLM applications, RAG architectures, agentic frameworks,
computer vision, NLP, and complete ML delivery.
- Lead the engineering of enterprise GenAI capabilities, including LLM orchestration, retrieval augmented generation, multi agent systems, and intelligent automation, across Rakuten Symphony product teams.
- AIML Platform & Infrastructure Engineering • Own the design and implementation of major platform subsystems: centralised inferencing stack, model registry, feature store integrations, and MLOps tooling.
- Lead the engineering of enterprise grade CI/CD pipelines for ML, covering training, evaluation,
deployment, canary rollouts, and automated rollback.
- Design and implement comprehensive model observability, including performance monitoring, data and concept drift detection, prediction quality tracking, and alerting.
- AI Quality Engineering and Evaluation • Own the organisation's approach to AI/ML quality assurance, defining evaluation methodologies,
benchmark datasets, and automated quality gates that govern promotion from development to production.
- Design and implement evaluation harnesses for LLM based systems: automated evaluation pipelines
(RAGAS or equivalent), human evaluation workflows, adversarial test suites, and regression benchmarks.
- AI Governance, Explainability and Responsible AI • Own the engineering implementation of AI governance across the platform: model risk management frameworks, audit trails, data and model lineage tracking, and regulatory compliance covering data residency, GDPR, and applicable industry standards.
- Design explainability as an engineering discipline, not as an afterthought. Build explanation APIs,
SHAP/LIME integration into serving pipelines, model cards, and automated bias detection into the standard model release workflow.
- AI Security Engineering • Own AI specific security threat modelling and mitigations, including prompt injection,
jailbreaking, data poisoning, model extraction, adversarial inputs, and PII leakage through generative outputs.
6.
Software
Architecture & Engineering Standards • Translate architectural blueprints into detailed technical designs, engineering specifications, and working implementations, not just documentation.
- Performance, Scalability & Reliability Engineering • Lead performance engineering for the inferencing stack: profiling, bottleneck identification, latency optimisation, and throughput scaling.
- Design for scale: thousands of concurrent inference requests, hundreds of models, and petabyte scale data volumes.
- Reusability, Inner-Sourcing & Developer Experience • Own the shared AI/ML services and libraries strategy, ensuring platform APIs are intuitive, clearly documented, and genuinely useful to product engineering teams.
- Drive inner sourcing initiatives so AI/ML assets, including models, pipelines, evaluation harnesses,
governance templates, and prompts, are discoverable, reusable, and governed across business units.
9.
Technical
Leadership & Mentorship • Provide senior technical mentorship to engineers across the AIML Framework team and product AI/ML teams, including on AI quality, governance, and responsible AI practices. REQUIRED KNOWLEDGE, SKILLS AND EXPERIENCE: Experience and Expertise
- 8 or more years of software engineering experience, with at least 5 years building production AI/ML systems spanning applied AI/ML delivery, platform infrastructure, and AI quality or governance.
- Demonstrated ability to design and build complete AI/ML systems from end to end, covering data and model development through production infrastructure,
evaluation, serving, and governance.
- Proven ownership of production GenAI or LLM systems: RAG architectures, fine tuned models, or agentic frameworks deployed at enterprise scale.
- Proven ownership of major AI/ML platform subsystems used by multiple teams,
such as inferencing services, model registries, and feature stores.
- Demonstrated experience implementing AI quality frameworks, evaluation pipelines, or model governance practices in a production environment.
- Track record of influencing engineering direction and raising engineering quality across teams without direct authority.
- Deep experience applying software engineering fundamentals to AI/ML systems: distributed systems, API design, system reliability, and production operations.
Technical Skills
Area 1 — Platform & Infrastructure o Expert knowledge of MLOps tools: MLflow, Kubeflow,
and at least one cloud native equivalent such as Vertex AI, SageMaker, or Azure ML.
o Expert level containerisation (Docker) and Kubernetes, including resource management and GPU scheduling for ML workloads.
o Solid data engineering skills: Spark, Kafka, data lakes, and feature stores.
o API design expertise for high availability, low latency inference services using
REST and gRPC.
o Experience with model optimisation techniques: quantisation, distillation,
ONNX, and TensorRT.
o Data and model lineage tooling to track provenance, transformations, and model to data traceability at enterprise scale. Area 2 — Applied AI/ML Engineering o Deep hands on expertise with TensorFlow and/or PyTorch, including model training, fine tuning, and deployment patterns.
o Production experience with LLM systems: prompt engineering, RAG pipeline design, retrieval optimisation, and structured output enforcement.
o Experience with agentic AI frameworks such as LangGraph, LangChain, or ADK for building multi agent production systems.
o Complete ML pipeline experience: feature engineering, model training,
evaluation, deployment, and monitoring in production.
o Experience with GenAI applications in enterprise contexts: document intelligence, knowledge retrieval, automation, or decision support.
o Breadth across AI modalities: NLP, computer vision, speech/audio, or time series, with production experience in at least two domains. Area 3 — Software Engineering Fundamentals o Expert level Python; proficiency in Go or Java strongly preferred.
o Strong software architecture and system design skills, covering distributed systems, microservices, event driven architectures, and API contracts.
o Proficiency in data structures, algorithms, and software engineering principles as applied to AI/ML system design.
o Experience with production software delivery practices: CI/CD, testing strategy
(unit, integration, and system tests), version control, and code review.
o Ability to write clean, maintainable, well tested code that other engineers can operate and extend.
Educational
Background • Bachelor's degree in Computer Science, Artificial Intelligence, Machine Learning, or related technical field. Master's preferred. RAKUTEN SHUGI PRINCIPLES:
Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.
- Always improve, always advance. Only be satisfied with complete success - Kaizen.
- Be passionately professional. Take an uncompromising approach to your work and be determined to be the best.
- Hypothesize - Practice - Validate - Shikumika. Use the Rakuten Cycle to success in unknown territory.
- Maximize Customer Satisfaction. The greatest satisfaction for workers in a service industry is to see their customers smile.
- Speed!! Speed!! Speed!! Always be conscious of time. Take charge, set clear goals, and engage your team.
📌 Principal Engineer – AI/ML (Bengaluru)
🏢 Rakuten Symphony
📍 Bengaluru