08 Sep
|
GuardianLink
|
Chennai
08 Sep
GuardianLink
Chennai
AI Engineer
What You’ll Do
Fine-tune and evaluate LLMs, including Indic-language models, using LoRA/QLoRA for lightweight tuning and full fine-tuning/RLHF for larger models.
Design and run multi-GPU and multi-node training jobs using H100/H200/B200-class hardware and InfiniBand-connected clusters.
Make cost-vs-reliability decisions on GPU procurement, including when to use spot instances with proper checkpointing versus on-demand instances for production-critical workloads.
Build and maintain inference pipelines, including RAG and agentic workflows, with a focus on latency and cost per request.
Containerize and deploy models using Docker and Kubernetes across cloud GPU providers.
Evaluate and benchmark GPU cloud options, including domestic and global neoclouds, based on cost, latency, and compliance fit.
Keep data-handling practices aligned with data residency and compliance requirements, including the DPDP Act, for regulated or personal-data workloads.
What You Should Know
Solid knowledge of ML frameworks including PyTorch, CUDA, and cuDNN.
Hands-on experience with LoRA,
QLoRA, RLHF, and distributed/multi-node training.
Experience with Docker, Kubernetes, and InfiniBand-connected GPU clusters.
Familiarity with H100, H200, B200, and A100 GPUs, including SXM and PCIe form factors.
Hands-on experience with at least one GPU cloud provider such as Spheron, Runpod, Vast.ai, E2E Networks, Lambda Labs, CoreWeave, or a similar provider.
Understanding of spot vs. on-demand trade-offs, checkpointing strategies for interruptible workloads, and per-minute/per-second billing optimization.
Awareness of data residency considerations under the DPDP Act 2023 for workloads involving personal data.
You Might Be a Fit If You Have
2+ years of experience building and shipping ML/AI systems in production.
Real-world experience training or fine-tuning transformer-based models, rather than just calling an API.
Experience working directly with GPU cloud infrastructure, including provisioni
📌 AI Engineer (Chennai)
🏢 GuardianLink
📍 Chennai