15 Aug
|
Hyper Lychee Labs
|
India
15 Aug
Hyper Lychee Labs
India
ENGAGEMENT: Full-time
WORK MODE: India (Remote / Off-site – quarterly in office visits to client site in Saudi Arabia)
JOB DESCRIPTION:
Key Responsibilities
Own the operationalisation of both traditional ML models and generative AI / LLM systems in a fully on-premise environment.
Deploy, monitor, and maintain LLM and agentic AI infrastructure running on client-owned GPUs.
Manage traditional MLOps responsibilities (model deployment, versioning, monitoring) alongside newer LLMOps workflows – this is a blended role, not LLM-only.
Work closely with the AI Infrastructure Architect to implement designed architectures for LLM/GPU management, parallelization, and agent orchestration.
Manage GPU resource allocation and token/context handling for on-premise LLM deployments (distinct from cloud-based token economics).
Support integration of MLOps pipelines with agentic and generative AI use cases as they move from design to production.
Must-Have Qualifications
Practical experience with LLM/generative AI deployment and management
Working knowledge of on-premise infrastructure management – this is not a cloud-native role. Candidates should understand how GPUs are provisioned, shared, and monitored outside of AWS/Azure/GCP-managed environments.
Understanding of how LLMs are stored and served across GPUs, and of parallelization approaches for splitting models/workloads across GPU resources.
Comfort working as part of a broader AI infrastructure team that includes an architect and, eventually, agent-orchestration workflows.
Secondary qualifications: Support MLOPs teams when necessary. Hands-on experience operationalizing ML models in production (traditional MLOps).
Nice-to-Have
Exposure to banking, financial services, or other highly regulated environments.
Familiarity with agentic AI frameworks and orchestration patterns.
On-premise GPU cluster management experience at scale.
Skills Required
Python, TensorFlow/PyTorch, GPU management (CUDA, NVIDIA tooling).
CI/CD for ML, Docker
📌 LLMOps Engineer (India)
🏢 Hyper Lychee Labs
📍 India