Sr Consultant
Noida, Uttar PradeshChennai, Tamil Nadu
Job Summary
Hands-on AI Consultant to design and operationalize domain SLM, RAG / GraphRAG, and multi-turn agentic solutions—from enterprise data ingestion and fine-tuning to low-latency serving, APIs, security, observability, HA/DR, and governance.
Key Responsibilities
Design enterprise GenAI platforms and domain SLMs using Gemma, Llama, Qwen, Phi, Mistral . Build agentic workflows for triage, diagnosis, remediation, approvals, escalation, and ticket automation. Develop ingestion pipelines for PDF, DOCX, PPTX, XLSX, SharePoint, APIs, and email , including OCR, deduplication, ACL preservation, lineage, monitoring, and PII controls. Lead SFT, LoRA/QLoRA, DPO/RLHF, distillation, quantization, and model compression . Architect RAG/GraphRAG using Qdrant, Neo4j, AuraDB, or Neptune with hybrid search, reranking, citations, and grounded retrieval. Deploy and optimize inference with vLLM, TensorRT-LLM, Triton, TGI, or SGLang , including batching, caching, streaming, autoscaling, and model routing. Build OpenAI-compatible APIs with tool calling, authentication, RBAC, rate limiting, and audit logging. Manage production deployments using Docker, Kubernetes, Helm, Terraform, CI/CD across Azure, AWS, and GCP. Define evaluation metrics for quality, grounding, hallucination, safety, latency, throughput, availability, and cost. Implement guardrails, observability, OpenTelemetry tracing, drift detection, governance, and compliance controls.
Skill Requirements
Core Area
Key Technologies & Techniques
Model Engineering
Python, PyTorch, Transformers, Hugging Face, TRL/PEFT, SFT, LoRA/QLoRA/DoRA, DPO/RLHF, BF16/FP16,
DDP/FSDP/ DeepSpeed ZeRO, MLflow / W&B.;
RAG & GraphRAG
Chunking, embeddings, hybrid search, reranking, Qdrant, Neo4j/AuraDB, Amazon Neptune, etc.; ontology, entity resolution, Cypher/Gremlin, vector-graph retrieval, grounding and citations.
Inference & APIs
vLLM, TensorRT-LLM, SGLang, TGI/Triton, batching, KV/prefix cache, speculative decoding, quantization, FastAPI / OpenAI-compatible APIs, structured output and tool calling; intelligent model routing using rules, semantic/complexity classifiers, cascades, cost-quality-latency policies, FinOps budgets, metering, fallback, and routing observability.
Platform & Resilience
Docker, Kubernetes, Helm, Terraform, CI/CD, Azure/AWS/GCP, multi-zone HA, cross-region DR, autoscaling, failover, backup/PITR, observability, RTO/RPO.
Security & Quality
Layered guardrails, prompt-injection defense, PII/secrets, RBAC/ACL, HITL, auditability, OpenTelemetry tracing, evaluation, drift monitoring, latency/throughput/cost SLOs.
Other Requirements
Experience & Qualifications
- 8-10 years in software, data, platform, or AI engineering; 5+ years in AI/ML and 4+ years in GenAI, SLM, or RAG.
- Hands-on ownership of a production domain model or enterprise RAG platform, with solid architecture, stakeholder, and cross-functional leadership.
- Bachelor’s or Master’s degree in Computer science / AI&ML; / Data Science / Engineering, or related discipline.
Preferred Candidate Profile
Hands-on architect who combines deep model, retrieval, platform, security, and operations expertise to move enterprise GenAI from experimentation to governed, highly available production.
📌 Sr Consultant (India)
🏢 HCLTech
📍 India