08 Sep
|
GuardianLink
|
Chennai
08 Sep
GuardianLink
Chennai
AI Engineer
What You’ll Do
- Fine-tune and evaluate LLMs, including Indic-language models, using LoRA/QLoRA for lightweight tuning and full fine-tuning/RLHF for larger models.
- Design and run multi-GPU and multi-node training jobs using H100/H200/B200-class hardware and InfiniBand-connected clusters.
- Make cost-vs-reliability decisions on GPU procurement, including when to use spot instances with proper checkpointing versus on-demand instances for production-critical workloads.
- Build and maintain inference pipelines, including RAG and agentic workflows, with a focus on latency and cost per request.
- Containerize and deploy models using Docker and Kubernetes across cloud GPU providers.
- Evaluate and benchmark GPU cloud options, including domestic and global neoclouds, based on cost, latency, and compliance fit.
- Keep data-handling practices aligned with data residency and compliance requirements, including the DPDP Act, for regulated or personal-data workloads.
What You Should Know
- Strong knowledge of ML frameworks including PyTorch, CUDA, and cuDNN.
- Hands-on experience with LoRA, QLoRA, RLHF, and distributed/multi-node training.
- Experience with Docker, Kubernetes, and InfiniBand-connected GPU clusters.
- Familiarity with H100, H200, B200, and A100 GPUs, including SXM and PCIe form factors.
- Hands-on experience with at least one GPU cloud provider such as Spheron, Runpod, Vast.ai, E2E Networks, Lambda Labs, CoreWeave, or a similar provider.
- Understanding of spot vs. on-demand trade-offs, checkpointing strategies for interruptible workloads, and per-minute/per-second billing optimization.
- Awareness of data residency considerations under the DPDP Act 2023 for workloads involving personal data.
You Might Be a Fit If You Have
- 2+ years of experience building and shipping ML/AI systems in production.
- Real-world experience training or fine-tuning transformer-based models, rather than just calling an API.
- Experience working directly with GPU cloud infrastructure, including provisioning, monitoring, and optimizing spend.
- Strong Python engineering fundamentals and experience with distributed systems.
- A bias toward shipping — you’d rather have a working V1 than a perfect plan.
Nice to Have
- Experience with Indic-language models or multilingual NLP.
- Exposure to regulated-industry inference pipelines, particularly in fintech or healthtech.
- Contributions to open-source ML tooling.
We are open to freelancers looking to transition into a full time opportunity as well as immediate joiners who can start at short notice.
📌 AI Engineer (Chennai)
🏢 GuardianLink
📍 Chennai