We are seeking an experienced AI developer to lead the fine-tuning, deployment, and optimisation of the custom Proniti AI model based on the prevailing AI model architecture (26B-A4B MoE / 31B Dense). You will be responsible for transforming the base model into a highly secure, autonomous reasoning engine capable of executing complex standard operating procedure (SOP) gap analyses and regulatory reporting.
Responsibilities- Model Fine-Tuning: Configure and execute supervised fine-tuning (SFT) pipelines using parameter-productive fine-tuning (PEFT) methodologies. Utilise Quantised Low-Rank Adaptation (QLoRA) with frameworks like Hugging Face TRL and Unsloth (using bitsandbytes NF4 quantisation) to adapt the model without catastrophic forgetting.
- Sovereign Infrastructure Deployment: Manage the deployment of the model on sovereign Indian cloud infrastructure. Work directly with dedicated infrastructure, NVIDIA H100 or L40S GPU clusters hosted in Mumbai-based Tier IV data centres to ensure data privacy and ultra-low latency.
- Inference Optimisation: Deploy and configure the vLLM inference engine. You will optimise the server using flags like '--gpu-memory-utilisation' for long context management and enable Gemma 4's specific parsers (--reasoning-parser gemma4). Agentic Tool Orchestration:
Implement native tool-calling capabilities by mapping Proniti's backend APIs to the model's < |tool_call|> and control tokens, powering the autonomous reporting agent.
- Constrained Decoding: Implement structured JSON output generation via vLLM's guided decoding engine to guarantee that the AI generates perfectly structured data payloads for the Proniti Compliance Dashboard.
- Security Governance: Integrate the open-source Agent Governance Toolkit to provide deterministic, sub-millisecond policy enforcement, preventing risks like tool misuse or prompt injections. Requirements
- 3-5+ years of experience in deep learning, NLP, and AI systems engineering.
- Strong proficiency in Python, PyTorch, and the Hugging Face ecosystem.
- Proven hands-on experience with LLM/SLM fine-tuning techniques (LoRA, QLoRA) and quantisation.
- Deep understanding of inference servers (specifically vLLM) and GPU memory optimisation (KV caching, PagedAttention).
- Experience building autonomous AI agents and utilising JSON schemas for strict output decoding.
- Qualifications: BE IT / BSc IT / MSc IT / MCA / ME IT / M. Tech IT or equivalent.
This job was posted by Sonali Betkar from AXS Solutions.
📌 AI Developer (Mumbai)
🏢 Axs Solutions
📍 Mumbai
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.